A trustworthy state-of-health number for every battery.
Physics-informed diagnostics from charge telemetry. No teardown, no test rig, no hardware on the vehicle — Cellworth reads the sessions your packs already log and issues a certified figure with the interval attached.
Validated, not assertedState-of-health error of 0.94–1.31 pp RMSE across 136 held-out cells in 3 public degradation datasets — NASA, Oxford, Stanford.
Read the protocolSpecimen certificate · CW-0417-A · issued from 540 EFC of telemetry
- State of health
- 87.4%
- 95% interval
- ± 1.8 pp
- Remaining life
- 630 EFC
- Evidence
- 61 charge sessions
- Hardware fitted
- None
Nobody can transact what nobody can measure.
sits in the pack — and it is the least knowable fact about the vehicle. Everything else on a used EV can be inspected in an afternoon.
Residual values, warranty reserves, insurance premiums and second-life prices all reduce to the same question: how much of this pack is left?
Today the answer comes from odometer proxies, a dashboard estimate nobody audits, or a teardown that costs more than the answer is worth.
The consequence is not that people get the number wrong. It is that whole transactions do not happen. A trader will not buy a pallet of packs they cannot grade. An insurer will not write a warranty against a distribution they cannot see. A fleet writes the asset down to the worst case because the worst case is the only case they can prove.
Telemetry in. A number you can defend out.
Three stages, and only the middle one is ours to keep. This page describes what Cellworth consumes and what it produces — the method itself is the asset.
Charging sessions, as they already exist
Pack voltage, current and temperature through a charge, at anything from 0.1 Hz upward, plus whatever state of charge the BMS already reports. Nothing is added to the vehicle and nothing is taken off the road.
- BMS and CAN logs
- OCPP charger session records
- Telematics and fleet-platform exports
An electrochemical model, fitted with machine learning
The degradation modes are written down as physics — lithium inventory loss, active-material loss, resistance growth — and the model is only allowed to move within them. Learning fills in what the physics leaves free.
- Extrapolates past the window it has seen
- Degrades gracefully on sparse, real-world data
- Every estimate decomposes into named modes
A figure with its interval attached
State of health and remaining useful life, each with a calibrated 95% interval and a breakdown of what is driving the fade. Delivered as a one-page certificate, a JSON payload, or a REST endpoint.
- Signed PDF certificate
- JSON payload per pack
- REST endpoint for portfolios
What “physics-informed” buys you over a score.
A black-box model trained on cycled cells learns the lab. Point it at a real vehicle — partial charges, a cold winter, a driver who never goes below 40% — and it has no basis for the extrapolation you are asking it to make, and no way to tell you it is guessing.
Constraining the model to known degradation physics changes what happens off the edge of the training data. The fade has to follow a shape the chemistry permits, the interval widens honestly as the evidence thins, and the output can be read back as a cause rather than a verdict: this pack has lost lithium inventory, not active material, which is why it looks the way it does.
The interval is the product. A number without one is a guess with better typography — and nobody underwrites a guess.
Which is why calibration is tested the boring way: on held-out cells, the stated 95% interval has to contain the true value 95% of the time. If it does not, the interval is wrong and the model does not ship.
Measured, not asserted.
Every figure below is measured on public degradation datasets anyone can download, against the methods a fleet or an insurer runs today, under a protocol stated in full. The ablations and the boundaries are here too — diligence should confirm this section, not dismantle it.
- Held-out cells · 3 datasets
- 0
- Pooled SoH RMSE
- 0.00pp
- Against the strongest baseline
- −0%
- Coverage of the 95% interval
- 0.0%
Results, per dataset
Error against the capacity the dataset's own reference test measured, on cells held out whole.
| Dataset | Cells | SoH RMSElower is better | SoH MAElower is better | p90 |error|lower is better | RUL MAPElower is better | vs. baselinebest published |
|---|---|---|---|---|---|---|
| NASA PCoENASA Ames Prognostics Center | 4 cells18650 LCO | 0.00pp | 0.00pp | 0.00pp | 0.0% | −0%1.38 pp baseline |
| Oxford degradationOxford Battery Intelligence Lab | 8 cellsKokam pouch NMC | 0.00pp | 0.00pp | 0.00pp | 0.0% | −0%1.44 pp baseline |
| Severson fast-chargeStanford · MIT · Toyota Research | 124 cellsA123 LFP / graphite | 0.00pp | 0.00pp | 0.00pp | 0.0% | −0%1.58 pp baseline |
Against the alternatives
An error figure alone is unreadable — 1.29 pp is excellent or useless depending on what the method you already run would have cost you. So here is that method, and the three between it and us.
Pooled SoH RMSE, cell-weighted across all 136 held-out cells. Every method was reproduced on the identical split — same cells, same charge segments, same reference capacities.
Is the interval honest?
The number a decision hangs on is the width of the band around it, so the band is tested on its own terms.
Dashed diagonal: a perfectly calibrated interval
| Stated interval | Observed coverage | Miss |
|---|---|---|
| 50% | 0.0% | +1.5 pp |
| 80% | 0.0% | −0.6 pp |
| 90% | 0.0% | −0.8 pp |
| 95% | 0.0% | −0.4 pp |
The band is the product, not the point estimate. A 95% interval that holds the truth 87% of the time is not cautious, it is wrong, and an underwriter who priced against it would find out two years later. Coverage is checked at four widths on cells the model has never seen; anything that drifts below its stated width is a defect, not a tuning choice.
What each part is worth
Remove a component, re-run the suite, report what breaks. The last row is the number this project would quote if it wanted to flatter itself.
| Configuration | Pooled RMSE | Δ | 95% coverage |
|---|---|---|---|
| Full modelPhysics-informed fade law, charge-segment features, per-cell interval calibration. | 1.29pp | — | 94.6% |
| − physics priorFree-form curve instead of the √n + linear fade law. Fits the observed window as well and extrapolates worse. | 1.61pp | +0.32 | 92.8% |
| − charge-segment featuresReference capacity tests only — the data a customer does not have. | 1.94pp | +0.65 | 93.5% |
| − interval calibrationSame point estimate, broken interval: the 95% band holds the truth 87% of the time. | 1.29pp | — | 87.1% |
| Random cycle splitleakageThe number you get by testing on cycles from cells the model trained on. Reported so you can recognise it elsewhere. | 0.62pp | −0.67 | 96.2% |
Protocol
Changing any of this changes the claim, so it is stated first.
- Split
- Held out whole cells, never held-out cycles. A cell seen in training never appears in test.
- Inputs
- Charge segments only — a partial constant-current window. No discharge capacity test, no reference performance test.
- Metric
- Error against the capacity measured by the dataset's own reference test, in percentage points of nameplate.
- Intervals
- Calibration checked by coverage: the stated 95% interval contains the true value on 95% of held-out cells, or the interval is wrong.
- Inference
- 38 ms per pack on one CPU core. No GPU anywhere in the loop.
- Training
- 11 minutes for the full suite on a laptop, from raw dataset files.
- Determinism
- Fixed seed, pinned environment. Every figure on this page regenerates from one command.
- Availability
- Evaluation code and split manifests go to counterparties and grant assessors on request.
Datasets: NASA Ames Prognostics Center of Excellence battery data; Oxford Battery Degradation Dataset 1; Severson et al. fast-charge dataset (Stanford, MIT, Toyota Research Institute). All three are public.
Where the evidence stops
Every figure above has an edge. We would rather draw it than have diligence find it — a limit you uncover yourself discredits everything standing next to it.
- Every cell in the table was cycled in a laboratory, at a temperature someone chose. Fleet telemetry is dirtier — dropped sessions, unknown ambient, a BMS that rounds. Nothing above proves the accuracy survives contact with it. The first pilot is what will.
- These are cells, not packs. Imbalance between cells is in the model and has not been checked against an instrumented pack, which is the widest gap between this table and a certificate issued on a vehicle.
- Coverage is LCO, NMC and LFP against graphite. Silicon-composite, LTO and sodium-ion stay out of scope until there is public data to hold them to: the fade law changes shape on those chemistries, and carrying this one across would be how a confident number becomes a wrong one.
- The intervals are calibrated on these duty cycles. Charge harder than the Severson envelope and the model is extrapolating — it widens the interval to say so, which is the correct behaviour and still not the same thing as evidence.
- None of this is a warranty-grade audit, and no quantity of public data would make it one. That takes a blind test on packs whose answer is already known — the first thing we ask a counterparty for, and the fastest way to prove or bury the claim.
Read the certificate before you talk to anyone.
One page, for a fictional pack, with the figure, the interval, the remaining-life curve and the methodology note. It is the whole product in a form you can forward to a colleague — which is the only test that matters.
Two markets, two artefacts.
Cellworth is not a battery analytics platform. It sells one object into each of two markets, and each object exists to unlock one decision.
A pilot is a file transfer.
The usual objection to battery analytics is that the integration is worse than the problem. It is worth being specific about how little Cellworth actually needs.
- Minimum input
- Eight charging sessions with at least 20 percentage points of state-of-charge swing. That is roughly a month of ordinary use, and it is enough to issue a certificate.
- Signals
- Pack voltage, pack current and at least one temperature, timestamped. State of charge and cell-level voltages improve the interval; neither is required.
- Sample rate
- 0.1 Hz and up. One reading every ten seconds through a charge is workable — most fleet platforms already log faster than that.
- Formats
- CSV or Parquet exports, OCPP 1.6 and 2.0.1 session records, raw CAN with a DBC, or a pull from your telematics API. We write the adapter.
- Not required
- No discharge capacity test, no reference performance test, no test rig, no workshop bay, no vehicle downtime, no hardware fitted to anything.
- First result
- Send an export on the Monday, get certificates back the same week. A pilot is a file transfer, not an installation project.
- Where it runs
- Processed and stored in the United Kingdom. No transfer outside the UK or EEA without a written agreement.
- Identifiers
- VINs and pack serials are pseudonymised on ingest and held separately from telemetry. We do not need to know whose vehicle it is to know how the pack is ageing.
- In transit and at rest
- TLS 1.3 in transit, AES-256 at rest, access scoped per engagement and logged.
- GDPR
- You are the controller, we are the processor. Standard DPA on request, deletion within 30 days of a written request, and a record of processing you can hand to your DPO.
- Your data stays yours
- Customer telemetry is not used to train models for anyone else unless you agree to it in writing, separately, and for a stated benefit.
Send telemetry. Get a certificate back.
The fastest way to find out whether this works on your packs is to let it fail on packs you already understand.
- A reply within two working days.
- First call is technical: what telemetry you hold, what decision you need the number for.
- A blind sample follows — you send packs you already know the answer on, and check us.
Or write to hello@cellworth.tech