The intelligence layer
A number is only useful if you know how wrong it can be.
Ionscore recovers a latent physical property from telemetry and reports it as a calibrated interval with a measured coverage rate. Two layers do the work: a degradation model with a physical functional form, fitted to real cells, and a learned layer that serves it in milliseconds. This page is both of them, stated plainly enough to be checked, including where they are weakest.
The measurement does not exist. It has to be inferred.
State of health is a latent physical property. You cannot read it off a sensor. The only honest way to obtain it directly is to take the pack out of the vehicle and run a full discharge cycle in a lab, which destroys the asset's availability for a day and costs more than the answer is worth.
So the problem is an inverse one: recover a hidden electrochemical state from the observable traces it leaves behind. Charge and discharge behaviour, resistance proxies, coulombic efficiency, thermal history, cycle and calendar ageing. This is what physical AI means in practice, and it is why a general-purpose model is the wrong tool. The structure of battery degradation is known physics; the estimator has to respect it.
The battery management system already reports a number, and that number is not reliable. A 2026 study across 1,114 vehicles from five manufacturers found BMS-reported state of health failing to track real capacity differences of 12 to 25 percent within the same model, with 384 of those vehicles not exposing the figure at all. An independent estimate is not a nicety. It is the only version of this number a third party can act on.
Physics underneath, learning on top.
A model trained only on telemetry learns correlations and inherits whatever bias the fleet it saw happened to carry. Ionscore is built the other way around: a degradation model with a physical functional form, fitted to real cells, and a learned layer that serves it fast.
Arrhenius, not curve fitting
Calendar fade follows the published Arrhenius x state-of-charge x square-root-of-time form; cycle fade follows charge throughput x depth of discharge x C-rate. Activation energies are physical parameters that get fitted, not free knobs.
A first-principles reference
A single-particle electrochemical model with an Arrhenius SEI-growth submodel, solved in PyBaMM, gives an independent physics reference. Where no field data exists, the semi-empirical model is checked against it for physical consistency.
Grounded, and held out
The coefficients are fitted to real NASA Ames PCoE cells, split cell-wise: fit on two cells, size a split-conformal band on a third, evaluate on a fourth the fit never saw. Disjoint cells, so the held-out claim is real.
The scope of that grounding is stated rather than implied: the NASA cells cover NMC cycle ageing at 24 C. LFP, calendar ageing, Indian heat and two- and three-wheeler duty cycles remain on literature defaults and are not yet validated against field data. Closing that gap is the research programme, and it is the part of this work that genuinely might not succeed.
Three models, then a proof about all three.
Ionscore does not fit one regressor and bolt an error bar onto it. It fits the interval directly, then earns the right to call that interval calibrated.
- 01
Quantile regression
Three gradient-boosted models are trained under the pinball loss, one each for the 10th, 50th and 90th percentile of state of health. The interval is a modelling target from the first epoch, not a post-hoc standard deviation.
- 02
Conformal calibration
The raw quantiles are then conformalised (CQR) on a calibration split the models never trained on. The residual distribution on that split sets how far each edge must move to hit its nominal rate.
- 03
Measured coverage
The widened interval is scored on a third, fully held-out split of 149 packs. Coverage comes back at 80.0 percent against an 80 percent target. That number is measured, and it is published whether or not it flatters us.
Why it matters
A split by pack, not by row. Every snapshot from a given battery lands on exactly one side of every split, so the model is never graded on a pack it has partly memorised. This is the difference between a coverage number that means something and one that measures leakage. The full card, including per-chemistry breakdowns and the conformal widening applied, is published on the methodology page.
Trained on synthetic fleets, and we will not bury that.
The model learns from simulated fleets grounded in public battery ageing data. That is a real limitation, it sets a ceiling we do not control, and it is the first question any competent reviewer should ask.
The published literature is clear about what happens next. A study from Bayreuth, the BMW Group and TUM reported the same architecture at three levels of realism: 1.98 percent RMSE in simulation, 2.95 percent on lab cells, and 8.56 percent in-vehicle. Our 1.58 percentage points of held-out error is structurally the first of those three numbers, and we expect field data to move it toward the third.
We are stating that in advance rather than discovering it in a pilot. The mitigation is also documented: a published result using 45 real-world segments alongside 1.7 million simulated ones reached 0.63 percent average error over a full lifespan. A small real anchor is worth more than a large synthetic corpus, which is exactly why the first design partner matters more to us than the next model revision.
Every competitor in this category publishes accuracy claims sourced only to themselves, and none of them publishes a calibration curve at all. We would rather be the one with a stated limitation and a measured coverage rate.
What we do not know yet.
Not shippedNothing in this section is in the product. These are the four open problems between the model as it stands and a model grounded on Indian field data, each written with the result that would tell us we were wrong.
- 01Open
Grounding on Indian duty and heat
The NASA cells ground NMC cycle ageing at 24 C. Indian two- and three-wheeler duty runs at 35 to 55 C with swap-station C-rates and deep discharge, and no open dataset covers it.
How we would know we are wrong
If fitted activation energies from Indian cells land far outside the literature range, the semi-empirical form does not transfer and the functional form itself has to change. - 02Open
Physics-informed estimation, not physics-generated data
Today the electrochemistry generates training data and audits the result. The stronger version constrains the estimator directly, so the degradation ODE is a term in the loss rather than a preprocessing step.
How we would know we are wrong
If the constrained estimator does not beat the unconstrained one on held-out real cells, the physics is decorative and we should say so and drop it. - 03Open
A small real anchor against a large synthetic corpus
Published work reached 0.63 percent average error using 45 real-world segments alongside 1.7 million simulated ones. Whether that ratio holds for 2W and 3W packs is untested.
How we would know we are wrong
If a few dozen real packs fail to move held-out error materially, the synthetic corpus is not the bottleneck and the whole data strategy needs rethinking. - 04Open
Weakest-cell inference from pack-level signals
A pack fails at its worst cell, but only pack-level telemetry is observable in the field. Recovering intra-pack spread from aggregate signals is the hardest open problem here.
How we would know we are wrong
If inferred cell spread does not correlate with teardown measurements on real packs, the signal is not recoverable at pack level and cell diagnostics stay a hardware problem.
All four need the same input, which is real packs under Indian conditions. That is why the first design partner is worth more to this company than the next model revision, and why the pilot is structured as a data exchange rather than a licence sale.
One inference pass, six decision surfaces.
A lender, an insurer, a recycler and an OEM warranty desk want different answers from the same pack. Scoring it once and projecting it six ways is cheaper and more consistent than four vendors disagreeing.
- State
Calibrated state of health
A P10 / P50 / P90 interval over usable capacity, plus a consumer-facing A to D grade. The interval is the product; the median alone would be a guess with better manners.
- Trajectory
Degradation projection
A forward curve from the current snapshot out to 60 months, so a five-year loan can be underwritten against the pack it will be, not the pack it is.
- Attribution
Reason codes
Every score decomposes into the drivers that moved it, each signed and ranked. A lender that cannot explain a decline cannot defend it.
- Behaviour
Usage and abuse signals
Thermal exposure, fast-charge fraction, depth-of-discharge habits and calendar ageing, separated from each other so a warranty exclusion can be argued from evidence.
- Value
Second-life routing
The projected crossing of the automotive-end threshold, and whether the pack is headed for repurpose, marginal or recycle, with the assumptions labelled.
- Exposure
Warranty loss cost
Breach probability against a contractual state-of-health floor and a monotone reserve curve, which is the number an OEM finance team actually carries.
Built to be consumed by agents.
An autonomous system cannot act on a bare number. It needs to know how much to trust it. A calibrated interval, a coverage guarantee and signed reason codes are precisely the inputs an agent needs to take a position it can later defend.
This engine was extracted from an autonomous underwriting product, so it was shaped for machine consumption before it ever had a user interface.
Shipped today
An autonomous covenant loop runs a lender's portfolio against its own policy on a caller-supplied as-of date, scores every position, and raises typed alerts for floor breaches, abuse signals and telemetry silence. It is deterministic by construction: the same inputs and the same as-of date always produce byte-identical output, which is what makes an automated decision auditable months later.
Designed, not yet built
A tool-callable surface over the same scoring contract, so an underwriting or remarketing agent can request a valuation, issue a certificate and read a portfolio without a human in the loop. The scoring API is already versioned, metered and schema-stable, which is the hard part. We would rather ship this with a design partner's real workflow than guess at one.
A model for the fleet that actually exists here.
India's electric fleet is two- and three-wheelers, and there is no open Indian battery dataset with state-of-health labels. Two independent curated inventories return zero Indian entries. A model tuned on European passenger cars is not transferable to a 2.8 kWh pack doing doorstep delivery in a coastal city at 38 degrees.
So the estimator is built around Indian duty cycles, chemistries and climate zones, and it runs entirely within Indian infrastructure. That matters commercially as well as technically: RBI's digital lending rules require borrower-linked data to stay on servers located in India.
The country is currently drafting the obligation before the measurement exists. The Battery Pack Aadhaar guidelines define health bands and name financiers and used vehicle buyers as stakeholders, while stating that the validation methodology is still to be developed. That is the gap this model is built to fill.
Point it at your packs.
A pilot scores a sample of your fleet end to end: calibrated state of health with published error bars, degradation, value, and risk, on your own telemetry.