Solar is unusually well instrumented and unusually badly analysed. Irradiance, module temperature, string current, inverter output and export metering are all measured from commissioning day. Most asset owners see a daily generation figure and a monthly performance ratio — one of which tells you nothing about where a problem is, and the other of which is a contractual metric rather than an operational one.
The gap between those and a genuinely actionable view is data aggregation done properly. This guide covers the decisions that determine whether that aggregation produces something an O&M team can act on, drawn from the plants and portfolios we monitor.
What you are actually aggregating
Four streams, each with different characteristics. Inverter and string telemetry: high volume, high cardinality, the source of nearly all loss attribution. Weather station data: low volume, high importance, because nothing can be normalised without plane-of-array irradiance and module temperature. Meter data: low volume, high authority, because this is what gets invoiced. And event data: alarms, trips, curtailment signals and maintenance records, which explain the gaps in everything else.
The most common mistake is treating these as one stream. They have different resolutions, different reliability and different roles. Meter data is the commercial truth and must reconcile exactly; inverter telemetry is diagnostic and can tolerate occasional gaps; weather data without gap-filling makes normalisation unreliable for the affected period, which must be marked rather than silently interpolated.
In practice
Resolution: the decision that constrains everything
Whatever resolution you store determines the analysis you will ever be able to do. Fifteen-minute data is adequate for reporting and monthly performance ratio. Five-minute is the practical minimum for meaningful loss attribution, because inverter clipping and derating behaviour smooth out at coarser intervals. One-minute is preferable for forecasting model training and for diagnosing intermittent faults.
Storage cost is rarely the constraint that people assume. A 50 MW plant with 1,200 inverters at one-minute resolution generates a substantial but entirely manageable volume in a columnar format with compression — the annual storage bill is trivial relative to a single day of lost generation.
What is a genuine constraint is what the SCADA historian retains and at what resolution it downsamples. Many systems keep high resolution for thirty days and then aggregate. If you intend to do serious analysis, extract at native resolution into your own store before the downsampling happens, because that history cannot be recovered later.
The retention trap
We have been called into several plants wanting two years of string-level analysis, only to find the historian downsampled to hourly after ninety days. Set up extraction to your own store on day one — the cost is negligible and the alternative is irreversible.
Normalising performance so comparison means something
Raw generation tells you almost nothing, because it varies with irradiance, temperature and time of year. Performance ratio normalises for irradiance; a temperature-corrected performance ratio also normalises for module temperature, which matters enormously in Indian summers where cell temperature routinely exceeds forty-five degrees.
Computed at string level rather than plant level, and compared across the plant, this becomes a map of where losses are concentrated. Present it as a ranked list rather than a chart: the fifty worst-performing strings this week, with an estimate of the generation each is losing and a probable cause category. That turns an analytical exercise into a work order, which is the entire point.
Because normalisation removes weather variation, a genuine degradation trend becomes visible early rather than being lost in the noise of a monsoon month — which is what allows a warranty claim to be made while the warranty is still live.
Separating soiling from shading from degradation
These three look similar in aggregate output and demand completely different responses. Soiling is recoverable by cleaning and follows a pattern driven by dust, pollen and rainfall. Shading is structural and time-of-day dependent. Degradation is permanent and should follow a slow, predictable curve.
Each leaves a distinct signature. Soiling produces a gradual decline across a whole area that recovers sharply after rain or cleaning. Shading produces a repeatable time-of-day pattern correlated with sun position and season. Degradation shows as a persistent offset that does not recover after cleaning.
The practical output is a cleaning decision supported by numbers rather than a calendar: the soiling loss on a given block has reached the point where cleaning cost is justified. Most plants clean on a fixed schedule. Plants that clean on measured soiling loss recover meaningfully more generation for the same expenditure — on one 120 MW portfolio, reallocating the same cleaning budget between blocks produced around two per cent additional generation.
Automating the DGR and PPA compliance pack
Every asset produces a daily generation report and at most sites somebody assembles it by hand from a SCADA export, a meter reading and the weather station. It takes an hour or two, it is inconsistent between sites, and it is late.
Automating it is the fastest return in the whole programme, because the labour saving starts in week one and it establishes the data flow that everything else depends on. Data read from plant systems, computed against your definitions, and published in your exact format to a shared location or emailed before the reporting deadline. For portfolios, every site producing an identically structured report is what finally makes fleet consolidation something other than a manual reworking exercise.
PPA compliance follows the same pattern, with one caution: contractual availability is rarely the same as raw uptime. Exclusions for grid unavailability, force majeure and scheduled maintenance must be implemented exactly as the contract defines them, and the evidence trail must survive an offtaker challenging the number.
Forecasting, and measuring whether it works
Day-ahead forecasting matters for scheduling and, in markets with deviation settlement, for avoiding penalties. Build it from numerical weather prediction inputs combined with the plant's own historical response, because two plants under identical forecast irradiance generate differently depending on configuration, soiling state and inverter behaviour.
Always present forecasts with an interval rather than a single number, and track accuracy continuously. A forecast whose historical error nobody measures is a guess with a chart attached. After a few months of learning, day-ahead accuracy within four per cent on a monthly aggregate basis is achievable and is materially better than the generic forecasts many operators rely on.
Key takeaways
- Four distinct streams — telemetry, weather, meters and events — with different roles and reliability requirements.
- Store at native resolution from day one; historian downsampling destroys history you cannot recover.
- Normalise for irradiance and temperature at string level, and present results as a ranked work list.
- Soiling, shading and degradation have distinct signatures and demand different responses.
- Automating the DGR pays back immediately and establishes the data flow everything else needs.
Frequently asked
No, and it is the normal case. Integration happens at the Modbus or SCADA level and the different register maps and naming conventions are normalised into a common model, so string-level analysis works identically across brands. Mapping takes a few days per manufacturer and only needs doing once.
Fifteen-minute for reporting and monthly PR. Five-minute or better for loss attribution and inverter behaviour analysis. One-minute is ideal for forecast model training. Start extracting at the highest resolution your SCADA exposes — you can always aggregate later, but you cannot recover resolution you never stored.
For a single site, often yes, and it removes any dependency on plant connectivity. For portfolios, cloud usually makes more sense because comparative analysis across sites is the point. Either way, plant-side collection should buffer locally so a connectivity drop delays rather than loses data.
Reporting automation saves labour from week one. Loss attribution typically produces its first actionable finding within six to eight weeks of baseline data, usually a set of underperforming strings or a soiling pattern that changes the cleaning schedule. Above roughly 20 MW the generation recovery alone generally covers the engagement within a year.