I want to… Understand the science
Science
How microbial electrochemical systems work, what they produce, and how performance is measured.
MESSAI · Onboarding · about 6 minutes
MESSAI connects microbial electrochemical research, structured experimental data, and modeling tools. Start here to understand the science, find your workspace, and see what the platform can — and cannot — support.
01
Four routes. Each one ends with something small and concrete you can actually do — the full list is in 08.
I want to… Understand the science
How microbial electrochemical systems work, what they produce, and how performance is measured.
I want to… Explore the research
Find papers, compare parameters, and inspect the evidence behind a result.
Open the research workspace Reference: Part II · 08 Pipeline →
I want to… Evaluate a process
Explore reactor configurations and economic assumptions — with their limitations visible.
I want to… Work on the platform
Get access, understand the architecture, and complete a small verified corpus run.
Engineer quickstart Reference: Part II · Running the corpus →
02
Every device in the taxonomy is a variation on one cell. What changes is the reaction you drive and the product you harvest.
Two facts govern everything downstream. The literature reports the same quantity in mutually incompatible ways — different normalisation bases, peak versus steady-state, undeclared reference electrodes — so raw numbers cannot simply be averaged. And the headline metrics are heavy-tailed across orders of magnitude, so a single number is almost always the wrong answer.
Full treatment: thermodynamics, kinetics, EET, the taxonomy →
03
A paper on its own is not data. It becomes usable when each measurement is attached to the conditions it was taken under, and normalised so it can sit beside a measurement from another paper.
Six stages, named once. The reference parts also describe a five-step acquisition sub-sequence, an eight-step runbook and a five-layer architecture view — those are substeps and layers of these six, and each says so where it appears.
04
One repository, four deployed zones, one PostgreSQL database behind them, and one object store for the PDF bytes. Only the API zone writes to the database.
messai-site
Marketing, learning and legal pages. Astro.
messai-ai
Research, papers, parameters, admin. The only zone with cross-zone rewrites.
messai-lab
The simulator, reactor models and 3D.
messai-api
Every route handler. The only zone that writes to the database.
The schema behind them carries 120 Prisma models, counted on
development on 2026-09-12.
The interactive architecture model, the schema browser and the package inventories →
05
Four different counting units, so these are cards and not a funnel: papers with extracted rows outnumber papers with a stored PDF, because legacy cohorts were imported without one. Each card names what it counts, where it was measured, and when.
Re-measure before quoting any of these. A number without its database and its date is not a number.
06
Deployment and validation are separate. A feature can be available and unvalidated; a number can be measured on staging and not yet deployed. The vocabulary below is used the same way everywhere on this page.
Search 23,579 papers and inspect where any extracted number came from
Every value carries its source paper; about half also carry the verbatim snippet, with the extractor version and a confidence. Where the snippet is missing, the record says so rather than implying a quote.
Get a calibrated interval for a performance parameter
The served hierarchical priors report a 90% band whose measured coverage is 89–91%. It is honest and it is wide — 4.1 decades on power density.
Beat a class-median baseline on paper-disjoint held-out data
No. On the one paper-disjoint held-out set, 0 of 4 targets pass the ≥10% RMSE-reduction gate against a class×domain median; typical held-out error on power density is 0.80 dex (about ×6) for every model and baseline alike.
Model an arbitrary paper end-to-end from its PDF
Harmonization is the binding constraint: 16% of extracted rows are modelable. Acquisition has also been stalled since the last discovery run in May 2026, so recent literature is thin — 23% of visible 2025–26 papers have a PDF.
Turn a curated paper into CapEx, OpEx, LCOE and a GWP offset
The TEA and LCA chain runs, and its assumptions are inspectable. Treat the output as a structured argument about which assumptions dominate, not as a costed quote.
Withdrawn claim: “Extraction succeeds on ~75% of papers”. Published with no cohort, denominator or date. Per-extractor row and paper counts are on the Atlas; a success RATE is not currently measured. What is measured instead →
07
These were argued and settled. The rejected options are recorded as history so the argument is not re-run — and, in the first case, so nobody picks up a fix that was explicitly ruled out.
Decided
Collapse the duplicates first, then add @@unique([paperId, label]), then emit sets at sync time and reuse an existing set instead of creating one.
Because
The unique key already names `runId`, which the backfill never sets; Postgres treats NULLs as distinct, so the constraint never fires.
Rejected — do not run
Setting runId on the backfill. It makes each re-run legitimately distinct — the exact behaviour being stopped. Do not run it.
Decided
Arm B — v2's flat transport with the v1.2 sub-prompt families ported onto it as separate flat calls. Measured on a 20-paper slice on 2026-09-10 (PR #902): 95% precision [92–98], 62% modelable, $0.05 per paper.
Because
v2's current prompt sourced 99 of 214 values from OTHER papers — reference lists and review tables — where arm B sourced 0 of 109. CMA v3 (Opus) recalls more but costs $0.93 per paper at 64% precision.
Rejected — do not run
v2's current prompt as-is (50% precision, and $0.16 per paper rather than the documented $0.05), CMA v3 for bulk, the three-pass variant, and resurrecting orchestrator.ts. Caveat that could reverse this: the gold set is LLM-adjudicated, not human-verified.
Decided
Three writer tiers, in this order: free regex over everything, then the title/abstract classifier ONLY where primarySystemType is still NULL, then paid full-text extraction last.
Because
Each tier costs more than the one before it, and each only touches what the cheaper tier could not settle, so the order is a cost order rather than a confidence order.
Rejected — do not run
Running the classifier first over the whole corpus. Documented exception to the tier order: classify-paper-type.ts nulls primarySystemType for REVIEW and FUNDAMENTAL papers as leakage, and both other writers put it back — so the result depends on which ran last.
08
Onboarding should end with the right next action, not with a finished reading list. Pick the one that matches your route.
09
Everything above is the summary. The detail lives on one page, the atlas, as five tabs — the full procedures, the commands and flags, the measured tables, and the incident history. Each subject has exactly one tab as its home.
docs/onboarding on GitHub — the markdown these parts were built from →