AI Solar Panel
011 AI Tools, Prompts and Data Pipelines 2,096 words · 10 min

Writing Prompts That Don’t Hallucinate Your Numbers

Last March I asked Claude to work out whether my 5.2 kWp array had covered the immersion heater run. I gave it the month’s generation total, the immersion’s rating, and the hours it had run. Back came a tidy answer: 312 kWh generated, 94 kWh to the immersion, 18% of output, £41 saved. Three of those four numbers were invented. My actual generation that month was 287 kWh. The model had rounded my figure up, then built a chain of arithmetic on the rounded value, then quoted a tariff rate I had never mentioned.

That is the core problem with using an LLM on your own solar data. It will not tell you when it has stopped reading your input and started filling gaps. The output looks identical either way: same confident tone, same clean pound signs. And because the numbers are plausible (312 is near 287, £41 is near what you’d guess), you don’t catch it by eye. You catch it six weeks later when your actual bill doesn’t match your model.

The fix is not a better model. It’s a better prompt, and specifically four patterns that each shut down one failure mode. I use these on Claude and ChatGPT for everything from payback modelling to reconciling Octopus Watch exports against my inverter logs. They work because they remove the model’s licence to guess.

Pattern one: paste the data, never describe it

The single biggest source of fabricated numbers is asking a question that references data the model cannot see.

Here is a prompt that will lie to you:

My 5.2 kWp array generated about 290 kWh in March.
My battery is 9.5 kWh. Work out how much of my
generation I self-consumed.

“About 290” gives the model permission to pick a number. It cannot see your consumption at all, so it will assume a UK household average, apply a typical self-consumption ratio (somewhere between 30% and 55% depending on which blog post is sitting in its weights), and present the result as your figure.

Here is the version that works:

Below is my March 2026 daily data, tab-separated.
Columns: date, generation_kwh, house_consumption_kwh,
grid_import_kwh, grid_export_kwh.

2026-03-01	4.21	9.84	6.90	1.27
2026-03-02	11.07	8.12	1.44	4.39
2026-03-03	9.88	10.51	3.02	2.39
...
2026-03-31	13.44	7.98	0.61	6.07

Use only these rows. Self-consumption =
generation - export. Give me the monthly total and
the ratio.

Now there is nothing to invent. Every input is on screen, and the definition of the thing you’re asking for is stated so the model isn’t choosing between three reasonable interpretations of “self-consumed”.

Getting the data out is easier than most people expect. GivEnergy, Solis and SolarEdge all expose CSV downloads in their portals; Octopus customers can pull half-hourly consumption via the Kraken API or the Octopus Watch app’s export; and if you’re on a smart meter with no useful app, your DCC data is available through a free service like Hugo Energy or n3rgy as half-hourly CSV going back a year. Thirty days of daily rows is roughly 1,200 tokens. You can paste that inline without thinking about it. A full year of half-hourly data (17,520 rows) needs a file upload and a code-execution tool rather than a chat paste, which is a different workflow and worth knowing the boundary of.

Pattern two: forbid estimation explicitly, and give it somewhere to put the gap

Telling a model “don’t estimate” is necessary but not sufficient. If you block the guess without providing an alternative, it will guess anyway, because answering feels more helpful than refusing.

The pattern that actually holds:

Rules for this task:
1. Use only numbers present in the data I pasted.
2. If a calculation needs a value I have not supplied,
   do not estimate it. Stop and list it under
   "MISSING INPUTS" with the units you need.
3. Do not substitute typical UK figures, averages, or
   values from your training data.

Rule 2 is the load-bearing one. It converts “I must produce an answer” into “I must produce an answer or a shopping list”, and the model will happily take the second option when it’s offered.

Watch what this catches. I asked for a payback calculation on adding 5 kWh of battery to an existing 9.5 kWh stack, and supplied generation, consumption, import and export. The response came back:

MISSING INPUTS
- Import tariff rate (p/kWh), and whether flat or
  time-of-use. If time-of-use, the rate bands and
  their hours.
- Export tariff rate (p/kWh) under your current SEG
  or Agile export arrangement.
- Standing charge (p/day) - needed only if you want
  total bill rather than marginal cost.
- Installed cost of the 5 kWh expansion including VAT.
- Round-trip efficiency of your battery, or I will
  need you to confirm I should treat it as lossless.

Without rule 2, that same request produced a complete payback figure of 7.4 years. It had assumed 24.5p import, 15p export, £2,400 install and 90% round-trip. Every one of those was wrong for me. My Agile import averaged 19.2p that month, my export is on the old 4.1p SEG rate, and the quote I had in hand was £1,850. Feeding the real numbers in gave 11.2 years. A 3.8-year error, dressed identically to the correct answer.

The round-trip efficiency line is the one I didn’t expect and now rely on. Battery losses are the classic silent assumption: the model treats charge and discharge as free, and your self-consumption number comes out 8-12% optimistic. Being asked about it made me go and find the actual figure from my GivEnergy spec sheet (86% AC-to-AC, not the 95% DC figure on the marketing page).

Pattern three: demand the arithmetic on screen

A stated answer cannot be audited. A shown calculation can, and the act of showing it changes what the model produces.

Add this:

Show every calculation as an explicit line:
  <input value> <operator> <input value> = <result>
Carry full precision through intermediate steps and
round only in the final answer, to 2 decimal places.
Do not state any number that is not either (a) from my
data or (b) the result of a shown line above it.

That last clause is the one worth copying verbatim. It creates a traceability requirement: every number in the output has to be either an input or the output of visible arithmetic. Fabrications have nowhere to hide, because a made-up number needs a made-up line above it, and made-up lines look wrong in a way made-up totals do not.

Here’s real output from that prompt, checking whether an October battery-charging strategy on Agile beat my previous flat rate:

Cheap-window import (00:30-04:30), 31 days:
  sum of import in window = 218.6 kWh
  218.6 x 0.0871 (avg Agile rate in window, £/kWh)
    = £19.04

Peak-window import (16:00-19:00), 31 days:
  sum of import in window = 47.3 kWh
  47.3 x 0.3412 = £16.14

All other hours:
  sum of import = 106.2 kWh
  106.2 x 0.1883 = £19.10

Total import cost = 19.04 + 16.14 + 19.10 = £54.28
Equivalent at flat 24.86p:
  372.1 x 0.2486 = £92.50
Saving = 92.50 - 54.28 = £38.22

I checked three of those lines on a calculator. All three held. More usefully, seeing the peak-window line made me notice that 47.3 kWh of peak import was far too high for a house with a 9.5 kWh battery that should have been full at 16:00. That turned out to be a real fault: my charge schedule was set to stop at 03:30, not 04:30, so the battery was reaching about 78% on winter nights. The shown arithmetic surfaced a hardware problem that a bare “you saved £38.22” would have buried.

Pattern four: make it state the method before it runs

The fourth pattern is the least obvious and prevents the subtlest error: the model doing correct arithmetic on the wrong model of your system.

Before calculating anything, state:
- The formula you will use, in symbols.
- Which column or value fills each symbol.
- Any assumption about system behaviour you are
  making (charge priority, export limits, losses).
Then wait for me to confirm before you calculate.

The “wait for me to confirm” turns it into a two-step exchange, which feels slower and saves far more time than it costs. When I applied this to a question about whether to charge the battery from grid overnight in shoulder months, the method came back as:

Formula: daily_cost = (import_cheap x rate_cheap) + (import_other x rate_other) - (export x rate_export)

Assumption: solar generation always charges the battery before exporting, and the battery never discharges to grid. Assumption: on days where overnight grid charge fills the battery, morning solar has nowhere to go and exports directly.

That second assumption is wrong for my setup. My GivEnergy Gen 3 is set to reserve 20% for backup, so it never fully fills from grid, and morning solar does have somewhere to go. Caught at the method stage it cost one sentence to fix. Caught after the calculation, I’d have had to work out which of eleven output lines were contaminated, and I probably wouldn’t have spotted it at all, because the final number would have looked sane.

This pattern also exposes silent unit mixing. Twice now a stated method has revealed the model planning to multiply kWh by p/kWh and report the result in pounds. Correct intent, wrong by a factor of 100, and completely invisible in a final figure that happens to land in a believable range.

Stacking them into a reusable block

The four patterns compose. I keep this as a saved prompt (a Claude Project custom instruction, though a text expander or a pinned note works the same) and paste it above any numeric question:

NUMERIC WORK RULES

DATA: use only values in the data block below. Never
substitute averages, typical UK figures, or values
from training data.

MISSING: if a value is needed and absent, stop and
list it under MISSING INPUTS with required units.
Do not proceed with a placeholder.

METHOD: state the formula in symbols, map each symbol
to a named column, and list every behavioural
assumption. Wait for my confirmation.

ARITHMETIC: after confirmation, show each step as
<value> <op> <value> = <result>. No number may appear
unless it is from my data or the result of a shown
line. Full precision through intermediates, round
only at the end.

UNITS: label every value with its unit. Flag any step
where units on the two sides do not reconcile.

Six hundred characters, and it changes the failure mode from “wrong answer that looks right” to “list of questions you can answer”. For a deeper treatment of how this fits with CSV exports, half-hourly reconciliation and the tooling around it, the AI tools, prompts and data pipelines pillar covers the plumbing side.

What the patterns cannot do

None of this makes the model good at arithmetic. On long multiplication and anything involving more than about six significant figures, token-by-token generation is genuinely unreliable, and no prompt fixes that. What the shown-arithmetic pattern gives you is not correctness, it’s checkability: you can verify 218.6 × 0.0871 in two seconds, and you cannot verify £54.28.

Where precision matters, push the calculation into code. Ask for Python rather than prose, and a model with a code interpreter will execute it and report real output:

Rather than calculating in text, write Python that
loads my pasted data and computes the result. Run it.
Show the code and its actual stdout. If you cannot
execute, say so and give me the code to run myself.

Pandas does not hallucinate a sum. The model can still write the wrong formula, which is exactly why pattern four sits upstream of this one: confirm the method, then let code do the multiplying.

One more habit worth building. Every time you get a numeric answer, pick one line at random and check it against source data. Not the total, a middle step. I do this maybe one time in five and I’ve caught two genuine errors in about forty sessions, both of them cases where the model had silently dropped rows (once a blank line in a pasted CSV, once a date that appeared twice because of a clock change on 29 March). Neither would have shown up in a plausibility check on the final figure. Both showed up immediately when I summed column two in LibreOffice and compared.