AI Solar Panel
§2 Section 2 of 6 2,627 words · 12 min

Forecasting What Your Roof Will Actually Generate

Every solar panel output forecast UK homeowners get from an installer lands somewhere between optimistic and fictional, and the gap is rarely malice. It’s that a single annual kWh number hides four or five separate models stacked on top of each other, each with its own error bars, and the quote only shows you the product. If you want a number you can actually plan a battery around, you have to take the stack apart.

This page is about doing that properly: picking a baseline irradiance dataset, building your own loss ladder instead of accepting a default, modelling shading that satellite data cannot see, converting an annual figure into the half-hourly shape that determines your bill, and then checking your model against your meter until the residuals stop telling you anything.

Three Different Forecasts, Not One

People say “output forecast” and mean three things that need different tools and have wildly different accuracy.

Annual energy yield is the long-run expectation: how much this roof produces in an average weather year. This is the number for payback calculations. Good datasets get you within 5% for an unshaded UK roof. It’s the easy one.

Monthly or seasonal shape matters because UK generation is brutally lopsided. A 5.2 kWp array in Sheffield producing 4,660 kWh a year does roughly 665 kWh in June and 85 kWh in December. That 7.8:1 ratio is why “my system covers 90% of my usage” is a statement about July, not about your annual bill.

Half-hourly production is what actually decides your economics, because self-consumption, battery cycling and export revenue all depend on whether a kilowatt-hour arrives at 13:00 or 18:30. Your annual model can be perfect and your savings estimate still off by 30% if the intraday shape is wrong.

Most homeowners model the first, assume the second, and never touch the third. The third is where the money is.

Pick a Baseline and Know Its Failure Modes

For UK roofs, PVGIS is the default starting point. It runs on satellite-derived irradiance (the SARAH-3 database covers Europe on roughly a 5 km grid, with hourly values back to 2005), it includes an SRTM-derived terrain horizon, and its API is free with no key.

A basic call looks like this:

https://re.jrc.ec.europa.eu/api/v5_3/PVcalc
  ?lat=53.383&lon=-1.466
  &peakpower=5.2&loss=14&mountingplace=building
  &angle=40&aspect=-20
  &raddatabase=PVGIS-SARAH3
  &usehorizon=1&outputformat=json

Two traps live in that URL. PVGIS measures azimuth with 0 = south, negative = east, positive = west, so aspect=-20 is a roof pointing 20° east of south (160° in compass terms). pvlib uses the opposite convention, 0 = north, 180 = south, and converts internally when it calls PVGIS. Get the sign wrong and your annual total barely moves, because east-of-south and west-of-south roofs at the same offset yield within about 1.5% of each other over a year. Your intraday curve, though, is mirrored, and your self-consumption estimate goes badly wrong in a way no annual sanity check will catch.

The other trap is loss=14. That’s a placeholder, not a model of your system, and we’ll replace it below.

PVWatts is the obvious alternative, and it is a genuinely good tool with a well-documented model chain, but its UK solar resource data comes from a coarser international dataset rather than the European satellite products PVGIS uses. The two disagree in ways that are systematic rather than random, especially in Scotland and on the western coasts. I’ve written up the comparison in detail, including side-by-side monthly outputs for four UK locations, at PVGIS vs PVWatts for UK Roofs. Short version: use PVGIS as your baseline, use PVWatts as a cross-check, and if they disagree by more than about 6% you have a siting or input error somewhere.

Worth adding a third opinion: Renewables.ninja gives hourly PV output for any UK coordinate using MERRA-2 reanalysis blended with CM-SAF satellite data, and it’s free for non-commercial use. Three independent estimates clustering within 5% is a much stronger signal than one estimate you trust.

Build the Loss Ladder Yourself

Set loss=0 in the PVGIS call and you get the array’s DC-side output before system losses. For our 5.2 kWp Sheffield array at 40° tilt, 160° azimuth, that comes out around 5,560 kWh. Now apply losses you can defend individually:

Loss termFactorReasoning
Inverter conversion0.970Weighted European efficiency from the datasheet, not peak efficiency
DC cabling + module mismatch0.980Standard for a two-string domestic install
AC cabling to the meter0.990Longer runs in detached properties; measure it
Soiling0.980Rain-washed at 40° pitch. Use 0.960 below about 12°
Snow0.997England and Wales. Scotland closer to 0.990
Year-one light-induced degradation0.985N-type TOPCon is better than this; PERC is worse
Shading0.935Measured, see below. Never guessed
Availability (faults, outages)0.990Roughly 3.5 days a year offline

Multiply through and you get 0.838, so 5,560 × 0.838 = 4,660 kWh in year one, or 896 kWh/kWp.

Notice what happened. The total loss is 16.2%, and PVGIS’s default was 14%. The default was close, but only because it’s a reasonable average for an unshaded array. PVGIS’s loss parameter contains no shading model at all beyond the terrain horizon, so accepting the default on a shaded roof means silently deleting your single largest loss term. That’s the most common way a domestic forecast ends up 10% high.

For subsequent years, apply linear degradation. Modern N-type modules warrant something like 1% in year one and 0.4% a year after that, reaching 87.5% at year 25. Year 10 output is about 4,480 kWh; the 25-year mean is roughly 4,240 kWh a year. Use the mean, not year one, when you compute payback.

Shading Is Where Forecasts Actually Break

PVGIS’s usehorizon=1 pulls a terrain horizon from SRTM elevation data at 3 arc-second resolution, about 90 m. It knows about the hill behind your village. It knows nothing about the neighbour’s chimney stack, the sycamore at the end of the garden, or your own dormer.

Measure the real horizon. Stand at the array’s centreline (or as close as you can get safely) and record the elevation angle of every obstruction by compass bearing. A phone clinometer app is accurate to a couple of degrees, which is plenty. Then feed it to PVGIS directly with userhorizon=, a comma-separated list of elevation angles starting due north and running clockwise at equal intervals.

Here’s why a few degrees matter. Suppose there’s a 9 m tree 12 m due south of an array mounted 3 m above ground. Elevation angle = atan((9 − 3) / 12) = 26.6°. Solar noon elevation at Sheffield’s latitude is 36.6° + declination. The sun clears the tree at noon only when declination exceeds −10°, which is roughly 6 March to 8 October. For five months of the year, that tree shades the array at the sunniest moment of the day.

Sounds catastrophic. It isn’t, because October through February together contribute only about 18% of annual yield in the UK. Model it properly and the tree costs maybe 6 to 8% annually. But it costs you 25 to 30% of your winter output, which is exactly when your heat pump wants it and when your battery arbitrage economics are tightest. An annual-only forecast tells you nothing about this.

Inverter topology changes the answer too. On a string inverter, a shadow crossing the bottom row of a 10-module string drags the whole string’s MPPT operating point down, and bypass diodes only partially rescue it. With per-module optimisers (SolarEdge P-series) or microinverters (Enphase IQ8), the loss stays roughly proportional to the shaded area. The difference between those two cases on a moderately shaded roof is commonly 8 to 15% of annual yield. pvlib doesn’t model partial shading well out of the box. NREL’s SAM does, via its 3D shade calculator plus the module-level shade loss database, and it’s free. If your roof has real obstructions, that hour spent in SAM is the highest-value modelling you’ll do.

Turning An Annual Number Into a Half-Hourly Profile

Grab hourly time series from PVGIS rather than the annual summary, because the annual summary throws away everything you need:

import pvlib

data, meta, inputs = pvlib.iotools.get_pvgis_hourly(
    latitude=53.383, longitude=-1.466,
    start=2016, end=2023,
    raddatabase='PVGIS-SARAH3',
    surface_tilt=40, surface_azimuth=160,   # pvlib convention: 180 = south
    pvcalculation=True, peakpower=5.2, loss=16.2,
    mountingplace='building', usehorizon=True,
    outputformat='json', map_variables=True,
)

That returns eight years of hourly output in UTC. Resample to half-hourly and you have the generation side. Now you need load at the same resolution, and this is the input nobody can synthesise for you.

If you’re with Octopus, their REST API gives half-hourly consumption straight from the DCC:

GET https://api.octopus.energy/v1/electricity-meter-points/{mpan}/meters/{serial}/consumption/
    ?period_from=2025-01-01T00:00Z&period_to=2026-01-01T00:00Z&page_size=25000

Other suppliers, use Hildebrand’s Bright app (free, and it exposes a Glowmarkt API) or n3rgy’s consumer service. One thing to check first: half-hourly data sharing is a consent setting, and plenty of SMETS2 meters are still set to daily. If yours is, change it and wait a month before you model anything, because daily totals are useless for this.

Then the arithmetic is trivial and the pitfalls aren’t:

self_consumption[t] = min(generation[t], load[t])
export[t]           = generation[t] - self_consumption[t]
import[t]           = load[t] - self_consumption[t]

PVGIS timestamps are UTC. Octopus returns UTC. Your smart meter app shows local time, which is UTC+1 from late March to late October. Join those on the wrong offset and you shift your entire summer generation curve an hour relative to load, which in a typical household moves modelled self-consumption by 5 to 8%. Every quantity in your model inherits that error.

One more correction to apply mentally: half-hourly averaging flatters self-consumption, because it smooths over the reality that a kettle draws 2.8 kW for four minutes while the array is doing 1.1 kW. Compared to 1-minute data, half-hourly typically overstates self-consumption by 2 to 3 percentage points, hourly by 5 or more. Shave your result accordingly rather than pretending it’s exact.

Pricing It: The Export Rate Dominates

Take our 4,660 kWh array against a 3,400 kWh/year household with a 250 W baseload. Running the half-hourly join gives about 1,350 kWh self-consumed with no battery, 29% of generation.

At 25.8p import and 15p export (Outgoing Octopus fixed):

Self-consumed  1,350 kWh × 25.8p = £348
Exported       3,310 kWh × 15.0p = £497
Annual value                     = £845

Add a 5 kWh usable battery and self-consumption rises to roughly 2,450 kWh. Recalculate:

Self-consumed  2,450 kWh × 25.8p = £632
Exported       2,210 kWh × 15.0p = £332
Annual value                     = £964

The battery earned £119 a year from solar shifting. Not £400. The reason is obvious once you see the arithmetic: the battery’s job is to convert kilowatt-hours from export value to import value, and the spread is only 10.8p, less round-trip losses.

Batteries pay in the UK mainly through tariff arbitrage, not solar. On Intelligent Octopus Go at 7p overnight against 24.5p day rate, cycling 5 kWh for 300 nights at 88% round-trip:

Displaced import   1,500 kWh × 24.5p =  £367.50
Overnight charging 1,705 kWh ×  7.0p = −£119.35
Net                                  =  £248

Those two streams don’t simply add. From May to August the battery is usually full from solar by mid-afternoon and has nothing left to buy cheaply, so the realistic combined figure is more like £300 to £340, against maybe £2,800 for 5 kWh installed. Eight to nine years, and that’s before you model degradation or a tariff change.

Flip the export rate and the whole design shifts. Those 3,310 exported kilowatt-hours are worth £497 at 15p and £136 at a 4.1p SEG rate. If you’re stuck on a low SEG, west-facing capacity and extra panels stop making sense and battery capacity starts making a lot more; on a good export tariff, the reverse. Model your tariff before you model your array.

For anything more sophisticated than “charge overnight, discharge on demand”, the two tools worth your time are Predbat and EMHASS, both Home Assistant add-ons. Predbat takes Solcast forecasts, your Octopus Agile rates and your own load history, and produces a half-hourly charge/discharge plan for the next 48 hours. EMHASS formulates it as a linear program via PuLP, which is more work to set up and more flexible if you want to add constraints like an EV or a heat pump.

Validate Against Your Meter, Then Read the Residuals

Twelve months in, compare. Here’s a real-shaped example against our 4,660 kWh model:

Month   Model   Meter    Error
Jan       130      96   -26.2%
Feb       220     178   -19.1%
Mar       370     358    -3.2%
Apr       545     585    +7.3%
May       645     690    +7.0%
Jun       665     640    -3.8%
Jul       635     612    -3.6%
Aug       550     572    +4.0%
Sep       420     402    -4.3%
Oct       275     231   -16.0%
Nov       120      88   -26.7%
Dec        85      61   -28.2%
------------------------------
Total   4,660   4,513    -3.2%

A 3.2% annual error looks like a good model. It isn’t. Look at the monthly column: summer within ±7%, winter consistently 16 to 28% low, with the error growing as the sun gets lower. That’s a shading signature, and specifically a low-elevation obstruction the model doesn’t know about. Random weather variation doesn’t produce monotone seasonal structure like that.

Two cautions when you read these numbers. Single-year UK irradiance varies about ±8% around the long-run mean, so one year’s total tells you very little on its own; the pattern of residuals is the informative part. And spring 2025 was the sunniest in the Met Office series, so the +7% in April and May is probably weather, not model error. Cross-check against Sheffield Solar’s PV_Live, which publishes half-hourly national PV output estimates. If PV_Live says the national fleet ran 9% above the seasonal norm in May and your array ran 7% above your model, your model is fine.

Track mean bias error month by month rather than a single annual number. When the monthly MBE loses its seasonal structure and sits inside roughly ±6% with no pattern, you’ve extracted what’s extractable and the rest is weather.

What AI Tooling Is Actually Good For Here

Ask Claude or GPT “what’s the annual yield of a 4 kWp south-facing array in Leeds” and you’ll get a confident number with no provenance. Don’t use it. Irradiance recall is exactly the task language models are worst at, and the answer will be plausible enough that you won’t question it.

Where they earn their keep is as glue. Handing a model the PVGIS API docs and asking it to write a fetcher that pulls eight years of hourly data, handles the UTC-to-Europe/London conversion correctly across DST boundaries, resamples to half-hourly and joins to an Octopus consumption export is a genuinely tedious hour of work that becomes ten minutes. So is asking it to convert a list of compass bearings and clinometer readings into a properly ordered userhorizon array. So is writing the LP formulation for battery dispatch against a 48-slot Agile price vector.

The discipline that makes this work: for anything numeric, make the model call a real tool rather than answer from memory, then verify the tool call independently. Run the same inputs through the PVGIS web interface and compare against your code’s output. If they differ by more than 0.5%, your code is wrong. That check takes two minutes and catches azimuth sign errors, unit confusion, and silently truncated date ranges.

Your next move, before any of the modelling above: get your consent set to half-hourly and start accumulating consumption data. Generation you can model from satellites. Your own load profile exists nowhere except your meter, and every number downstream of it is only as good as the year of history you haven’t started collecting yet.

In this section

The supporting pages under this subject.