← All decks

bes-benchmark-v1 · held-out prediction error · measured 2026-09-02

MESSAI · local corpus, read-only

bes-benchmark-v1

The first paper-disjoint, leakage-audited held-out set for bioelectrochemical-system performance prediction, and what every predictor MESSAI owns scores on it.

3,593 rows · 590 papers · 4 targetsbranch feat/bes-benchmark-v1
×6
typical held-out error on power density, for every model and baseline alike
0.80 dex
0 of 4
targets where any model passes the ≥10% RMSE-reduction gate against the class median
M1 · M2 · M4
96.7%
of benchmark rows the served physics predictor cannot score — it needs an HRT the paper never reported
3,463 of 3,593 rows
590
papers, three splits that pass the adversarial audit
hash · group 5-fold · temporal

The number MESSAI has been quoting since May — 372 held-out predictions, ECE 1.96% — matches no artifact. What does exist is weaker than it looks.

ArtifactWhat it saysWhat it actually is
calibration.json (live, 2026-07-22)517 rows · 95.9% coverage · ECE 0.036Random row-level 80/20 split of a 7-feature RandomForest. Papers straddle train and test. Its own power-density test R² is −3.7 × 10⁶. σ is re-estimated from the same held-out residuals, then pooled across volts, ohms and W/m², so the coverage is close to tautological.
“372 held-out, ECE 1.96%”quoted on 6 surfacesNo such file. The lab app bundled a third vintage (377 rows) at build time, so lab and web showed different numbers for the same concept.
fit_validation_role split97.98% OOS coverage · 940 obsNot exchangeable. A classifier on design covariates alone tells stored fit papers from validation papers at AUC 0.77–0.92 (permutation null ≈ 0.50). Every May coverage, conformal and PSIS-LOO number was computed on it.
Served 95% interval on /api/ml/predict"calibrated"A hard-coded ±25% band, multiplied by a conformal q̂ fitted to a different model's residuals. The version-drift warning fires on 100% of responses, correctly.

Three splits that pass the audit

Four targets from the modelable corpus, filtered in a recorded order. Two filters did most of the work: 1,128 rows were values a paper cited from another paper, and 1,372 were exact duplicates.

Filter funnel
rows remaining
all modelable rows
7,336
vintage v2 / curated
6,884
verifier not failed
6,466
not cited from another paper
5,338
no physics/dedupe flags
5,281
BES, not review
5,023
physical bounds
4,965
exact duplicates removed
3,593
Extractor vintage v2 or curated tuples only; verifier not failed; not cited from another paper; no physics or dedupe flags; not a review or non-BES paper; inside physical bounds; exact duplicates removed.
Split audit — adversarial AUC
paper-level · 0.5 = exchangeable
SplitpowercurrentCECOD
stored fit_validation_role0.8990.9200.7700.845
hash split0.5190.4700.4640.531
group 5-fold0.6000.6180.4120.602
temporal ≥ 20240.4290.5240.5800.491
A LightGBM classifier on design covariates, one vector per paper, 5-fold. Above 0.65 the split leaks. The stored split leaks on every target; the benchmark splits sit at the null.
TargetTransformRowsPapersWithin-paper SDBetween-paper SD
power_density_areallog10 W/m²1,0373050.61 dex1.10 dex
current_density_areallog10 A/m²9112420.61 dex1.16 dex
coulombic_efficiencylogit6001900.962.04
cod_removallogit1,0452910.861.35

Between-paper spread is roughly twice within-paper spread. There is cross-paper variance a covariate model could explain. The next section shows none of them do.

Median absolute held-out error, group 5-fold out-of-fold, with paper-grouped bootstrap 95% CIs. The green row is the bar: the median of the training papers in the same system class and application domain. B3 is the served physics predictor, scored only where it can predict.

power_density_areal
median |error| · dex
ModelMedian95% CI
B0 · global median0.800[0.679, 0.948]
B1 · class×domain median (the bar)0.795[0.645, 0.952]
B2 · served priors v20.831[0.710, 0.968]
M1 · LightGBM, 35 design covariates0.812[0.721, 0.940]
M2 · class median + residual GBM0.813[0.704, 0.940]
M4 · RidgeCV, one-hot + imputed0.942[0.797, 1.093]
B3 · served physics predictor, only where it can predict0.462[0.351, 1.659] · n=15/1,037
Bar: 0.80 dex ≈ ×6.2. Group 5-fold OOF, paper-bootstrap 95% CI.
current_density_areal
median |error| · dex
ModelMedian95% CI
B0 · global median0.885[0.741, 1.016]
B1 · class×domain median (the bar)0.784[0.666, 0.941]
B2 · served priors v20.719[0.644, 0.839]
M1 · LightGBM, 35 design covariates0.895[0.755, 1.025]
M2 · class median + residual GBM0.824[0.669, 1.022]
M4 · RidgeCV, one-hot + imputed1.140[0.962, 1.359]
B3 · served physics predictor, only where it can predictno rows scorable
Bar: 0.78 dex ≈ ×6.1. Group 5-fold OOF, paper-bootstrap 95% CI.
coulombic_efficiency
median |error| · logit
ModelMedian95% CI
B0 · global median1.474[1.299, 1.683]
B1 · class×domain median (the bar)1.194[1.003, 1.374]
B2 · served priors v21.199[1.020, 1.373]
M1 · LightGBM, 35 design covariates1.173[1.019, 1.331]
M2 · class median + residual GBM1.096[0.968, 1.320]
M4 · RidgeCV, one-hot + imputed1.317[1.190, 1.485]
B3 · served physics predictor, only where it can predict1.953[1.074, 3.189] · n=22/600
Bar: 1.19 logit. Group 5-fold OOF, paper-bootstrap 95% CI.
cod_removal
median |error| · logit
ModelMedian95% CI
B0 · global median0.936[0.849, 1.053]
B1 · class×domain median (the bar)1.056[0.939, 1.153]
B2 · served priors v20.989[0.863, 1.085]
M1 · LightGBM, 35 design covariates0.992[0.844, 1.128]
M2 · class median + residual GBM1.013[0.882, 1.170]
M4 · RidgeCV, one-hot + imputed1.078[0.951, 1.200]
B3 · served physics predictor, only where it can predict1.306[0.721, 1.908] · n=58/1,045
Bar: 1.06 logit. Group 5-fold OOF, paper-bootstrap 95% CI.
Gate · ≥10% RMSE cut vs B1, CIs disjoint, on both paper-disjoint splits, direction-correct on 2024+powercurrentCECOD
M1 · LightGBM, 35 design covariatesfail −3.2%fail −1.5%fail −4.0%fail +1.5%
M2 · class median + residual GBMfail −6.0%fail −4.0%fail −2.8%fail −0.9%
M4 · RidgeCV, one-hot + imputedfail −21%fail −33%fail −24%fail −8%
Same gate, one peak value per paperfailfailfailfail

Top GBM gain features are substrate concentration, publication year, reactor volume and electrode area at 9–18% each. No physical driver dominates, and the covariates the corpus holds for these rows are thin: temperature 49%, pH 44%, anode material 74%, HRT 6%, inoculum 0%.

90% interval coverage
target 0.90
BandpowercurrentCECOD
B2 · served priors v20.890.900.920.82
B1 · residual band0.880.850.870.86
M3 · conformal quantile regression0.910.920.900.87
served ±25% band0.00—0.000.19
The served ±25% band covers 0% of power-density rows and 19% of COD-removal rows where it predicts at all.
Interval width
decades (log10) · logit units for fractions
BandpowercurrentCECOD
B2 · served priors v24.14.26.65.8
B1 · residual band4.03.85.94.5
M3 · conformal quantile regression4.44.66.54.8
served ±25% band0.2—0.70.7
Coverage is bought with width: a 4-decade band on power density spans 0.0001 to 10 W/m². M3 is never narrower than the class-median band.
Why predictForSystem could not score the literature
rows · no fabricated inputs
predicted
118 · 3.3%
insufficient inputs
3,463 · 96.4%
unroutable class/unit
12 · 0.3%
missing-input combinations, rows (top 6 of 15)
HRT only
959
all four
482
pH + HRT
453
temperature + HRT
412
temperature + pH + HRT
344
HRT + COD
247
The predictor requires temperature, pH, HRT and COD. HRT alone is the only missing input on 959 rows; batch reactors do not have one. The old skill harness filled the gaps with 30 °C, pH 7, 12 h and 1000 mg/L. This dump does not.

Published R² of 0.95 to 0.997 for MFC power density come from within-lab random splits of one dataset. That is the artefact this platform already diagnosed in its own trainers: train R² 1.000, out-of-fold below zero. On a paper-disjoint benchmark the field's number, measured here for the first time, is a ×6 typical error and R² ≈ 0 for any covariate model.

Accuracy will not move with another model. It moves when complete design → conditions → outcome tuples exist for enough papers, which is a targeted re-extraction, not a fit.

Next
  1. Rewire the MFC interval: serve the priors' predictive band instead of σ = 25% of the point, and return insufficient_inputs instead of the 46-paper heuristic when design covariates are absent.
  2. Re-run run_benchmark_v1; the served band should then cover ~90% instead of 0%.
  3. Score the four expert datasets tagged holdout_set as an external test.
  4. Smoke a targeted re-extraction of full tuples on 20 papers before spending on the corpus.

The roadmap's model-quality workstream reports against this table. A model earns a row only by passing the gate above on bes-benchmark-v1; until one does, the baseline is the scoreboard. Median absolute error, group 5-fold out-of-fold.

ModelMetricHeld-out scoreDate
prior mean (baseline) · B1 class×domain medianmedian |error| · power_density_areal0.795 dex (≈ ×6.2) · CI [0.645, 0.952]2026-09-02
No model scored yet — B2-M1 (due 10 Oct 2026)≥10% RMSE cut vs B1——
What we do not claim
  • M1, M2 and M4 all fail the gate on all four targets (−33% to +1.5% RMSE change vs B1); they are not listed as scored models.
  • B3, the served physics predictor, scores 118 of 3,593 rows and is excluded until it can predict without fabricated inputs.
  • Scores here are the artifact’s 2026-09-02 numbers; re-run run_benchmark_v1 before adding a row.
Sources: bes-benchmark-v1 artifact (measured 2026-09-02, local corpus, read-only) · branch feat/bes-benchmark-v1 · services/ml-engine/training/benchmark/ · handoff docs/handoffs/2026-09-02-bes-benchmark-v1.md · 34 offline tests · leakage audit in manifest.json · calibration.json vintage 2026-07-22.