bes-benchmark-v1 · held-out prediction error · measured 2026-09-02
MESSAI · local corpus, read-only
bes-benchmark-v1
The first paper-disjoint, leakage-audited held-out set for bioelectrochemical-system performance prediction, and what every predictor MESSAI owns scores on it.
The number MESSAI has been quoting since May — 372 held-out predictions, ECE 1.96% — matches no artifact. What does exist is weaker than it looks.
| Artifact | What it says | What it actually is |
|---|---|---|
| calibration.json (live, 2026-07-22) | 517 rows · 95.9% coverage · ECE 0.036 | Random row-level 80/20 split of a 7-feature RandomForest. Papers straddle train and test. Its own power-density test R² is −3.7 × 10⁶. σ is re-estimated from the same held-out residuals, then pooled across volts, ohms and W/m², so the coverage is close to tautological. |
| “372 held-out, ECE 1.96%” | quoted on 6 surfaces | No such file. The lab app bundled a third vintage (377 rows) at build time, so lab and web showed different numbers for the same concept. |
| fit_validation_role split | 97.98% OOS coverage · 940 obs | Not exchangeable. A classifier on design covariates alone tells stored fit papers from validation papers at AUC 0.77–0.92 (permutation null ≈ 0.50). Every May coverage, conformal and PSIS-LOO number was computed on it. |
| Served 95% interval on /api/ml/predict | "calibrated" | A hard-coded ±25% band, multiplied by a conformal q̂ fitted to a different model's residuals. The version-drift warning fires on 100% of responses, correctly. |
Three splits that pass the audit
Four targets from the modelable corpus, filtered in a recorded order. Two filters did most of the work: 1,128 rows were values a paper cited from another paper, and 1,372 were exact duplicates.
| Split | power | current | CE | COD |
|---|---|---|---|---|
| stored fit_validation_role | 0.899 | 0.920 | 0.770 | 0.845 |
| hash split | 0.519 | 0.470 | 0.464 | 0.531 |
| group 5-fold | 0.600 | 0.618 | 0.412 | 0.602 |
| temporal ≥ 2024 | 0.429 | 0.524 | 0.580 | 0.491 |
| Target | Transform | Rows | Papers | Within-paper SD | Between-paper SD |
|---|---|---|---|---|---|
| power_density_areal | log10 W/m² | 1,037 | 305 | 0.61 dex | 1.10 dex |
| current_density_areal | log10 A/m² | 911 | 242 | 0.61 dex | 1.16 dex |
| coulombic_efficiency | logit | 600 | 190 | 0.96 | 2.04 |
| cod_removal | logit | 1,045 | 291 | 0.86 | 1.35 |
Between-paper spread is roughly twice within-paper spread. There is cross-paper variance a covariate model could explain. The next section shows none of them do.
Median absolute held-out error, group 5-fold out-of-fold, with paper-grouped bootstrap 95% CIs. The green row is the bar: the median of the training papers in the same system class and application domain. B3 is the served physics predictor, scored only where it can predict.
| Model | Median | 95% CI |
|---|---|---|
| B0 · global median | 0.800 | [0.679, 0.948] |
| B1 · class×domain median (the bar) | 0.795 | [0.645, 0.952] |
| B2 · served priors v2 | 0.831 | [0.710, 0.968] |
| M1 · LightGBM, 35 design covariates | 0.812 | [0.721, 0.940] |
| M2 · class median + residual GBM | 0.813 | [0.704, 0.940] |
| M4 · RidgeCV, one-hot + imputed | 0.942 | [0.797, 1.093] |
| B3 · served physics predictor, only where it can predict | 0.462 | [0.351, 1.659] · n=15/1,037 |
| Model | Median | 95% CI |
|---|---|---|
| B0 · global median | 0.885 | [0.741, 1.016] |
| B1 · class×domain median (the bar) | 0.784 | [0.666, 0.941] |
| B2 · served priors v2 | 0.719 | [0.644, 0.839] |
| M1 · LightGBM, 35 design covariates | 0.895 | [0.755, 1.025] |
| M2 · class median + residual GBM | 0.824 | [0.669, 1.022] |
| M4 · RidgeCV, one-hot + imputed | 1.140 | [0.962, 1.359] |
| B3 · served physics predictor, only where it can predict | no rows scorable |
| Model | Median | 95% CI |
|---|---|---|
| B0 · global median | 1.474 | [1.299, 1.683] |
| B1 · class×domain median (the bar) | 1.194 | [1.003, 1.374] |
| B2 · served priors v2 | 1.199 | [1.020, 1.373] |
| M1 · LightGBM, 35 design covariates | 1.173 | [1.019, 1.331] |
| M2 · class median + residual GBM | 1.096 | [0.968, 1.320] |
| M4 · RidgeCV, one-hot + imputed | 1.317 | [1.190, 1.485] |
| B3 · served physics predictor, only where it can predict | 1.953 | [1.074, 3.189] · n=22/600 |
| Model | Median | 95% CI |
|---|---|---|
| B0 · global median | 0.936 | [0.849, 1.053] |
| B1 · class×domain median (the bar) | 1.056 | [0.939, 1.153] |
| B2 · served priors v2 | 0.989 | [0.863, 1.085] |
| M1 · LightGBM, 35 design covariates | 0.992 | [0.844, 1.128] |
| M2 · class median + residual GBM | 1.013 | [0.882, 1.170] |
| M4 · RidgeCV, one-hot + imputed | 1.078 | [0.951, 1.200] |
| B3 · served physics predictor, only where it can predict | 1.306 | [0.721, 1.908] · n=58/1,045 |
| Gate · ≥10% RMSE cut vs B1, CIs disjoint, on both paper-disjoint splits, direction-correct on 2024+ | power | current | CE | COD |
|---|---|---|---|---|
| M1 · LightGBM, 35 design covariates | fail −3.2% | fail −1.5% | fail −4.0% | fail +1.5% |
| M2 · class median + residual GBM | fail −6.0% | fail −4.0% | fail −2.8% | fail −0.9% |
| M4 · RidgeCV, one-hot + imputed | fail −21% | fail −33% | fail −24% | fail −8% |
| Same gate, one peak value per paper | fail | fail | fail | fail |
Top GBM gain features are substrate concentration, publication year, reactor volume and electrode area at 9–18% each. No physical driver dominates, and the covariates the corpus holds for these rows are thin: temperature 49%, pH 44%, anode material 74%, HRT 6%, inoculum 0%.
| Band | power | current | CE | COD |
|---|---|---|---|---|
| B2 · served priors v2 | 0.89 | 0.90 | 0.92 | 0.82 |
| B1 · residual band | 0.88 | 0.85 | 0.87 | 0.86 |
| M3 · conformal quantile regression | 0.91 | 0.92 | 0.90 | 0.87 |
| served ±25% band | 0.00 | — | 0.00 | 0.19 |
| Band | power | current | CE | COD |
|---|---|---|---|---|
| B2 · served priors v2 | 4.1 | 4.2 | 6.6 | 5.8 |
| B1 · residual band | 4.0 | 3.8 | 5.9 | 4.5 |
| M3 · conformal quantile regression | 4.4 | 4.6 | 6.5 | 4.8 |
| served ±25% band | 0.2 | — | 0.7 | 0.7 |
Published R² of 0.95 to 0.997 for MFC power density come from within-lab random splits of one dataset. That is the artefact this platform already diagnosed in its own trainers: train R² 1.000, out-of-fold below zero. On a paper-disjoint benchmark the field's number, measured here for the first time, is a ×6 typical error and R² ≈ 0 for any covariate model.
Accuracy will not move with another model. It moves when complete design → conditions → outcome tuples exist for enough papers, which is a targeted re-extraction, not a fit.
- Rewire the MFC interval: serve the priors' predictive band instead of σ = 25% of the point, and return insufficient_inputs instead of the 46-paper heuristic when design covariates are absent.
- Re-run run_benchmark_v1; the served band should then cover ~90% instead of 0%.
- Score the four expert datasets tagged holdout_set as an external test.
- Smoke a targeted re-extraction of full tuples on 20 papers before spending on the corpus.
The roadmap's model-quality workstream reports against this table. A model earns a row only by passing the gate above on bes-benchmark-v1; until one does, the baseline is the scoreboard. Median absolute error, group 5-fold out-of-fold.
| Model | Metric | Held-out score | Date |
|---|---|---|---|
| prior mean (baseline) · B1 class×domain median | median |error| · power_density_areal | 0.795 dex (≈ ×6.2) · CI [0.645, 0.952] | 2026-09-02 |
| No model scored yet — B2-M1 (due 10 Oct 2026) | ≥10% RMSE cut vs B1 | — | — |
- M1, M2 and M4 all fail the gate on all four targets (−33% to +1.5% RMSE change vs B1); they are not listed as scored models.
- B3, the served physics predictor, scores 118 of 3,593 rows and is excluded until it can predict without fabricated inputs.
- Scores here are the artifact’s 2026-09-02 numbers; re-run run_benchmark_v1 before adding a row.