← All decks

Platform map · orientation · measured 2026-09-03

samfrons/messai.ai · feat/corpus-derived-recommendations

One repo, four Vercel zones, nine open-source packages, one Postgres.

Every surface MESSAI ships, what feeds it, and what is still dark. Requests enter apps/web and are rewritten to site, lab and api by URL path. Only apps/api writes to Postgres. Everything to the left of the zones is a person or a script; everything to the right is state or a model.

system mapdata pipelineopen sourceschema · 117 modelsknown gaps · 13
Branches
Drawn from feat/corpus-derived-recommendations, 11 commits ahead of development. Every count is identical on development except the GapSearchRun model and the effects-refit step. main (production) is 63 commits behind development, so prod lags this picture by the same amount.
Verified vs. quoted
Route, page, model, lib and workflow counts and the BullMQ mock were measured from the tree on 2026-09-03. Corpus row counts, R2 and HF Hub sizes, and the GP-SCM fit status come from dated docs and were not re-measured against a database.
4
Vercel zones
one repo
224
API routes
apps/api · measured 2026-09-03
100
pages
web 84 · lab 16
14
@messai/* shared libs
libs/

Batch work has no worker host: scripts run on an operator's machine or in the weekly Claude routine. The BullMQ box is dormant scaffold. Blue cards are Vercel zones, green are shared libraries or live services, amber are scripts, dashed are dormant.

Clients
Researcher browser
client
messai.io
HTTPS

Every request enters through the messai-ai Vercel project (apps/web), which owns / and rewrites everything else to the other three zones.

entry
apps/web/next.config.js
rewrite map
apps/web/multi-zone-rewrites.json
blocks
api 36 · lab 10 · site 17 entries
AI chat & agents
client
/api/chat · tool registry

One canonical RAG chat endpoint with the full tool registry (chat bifurcation resolved 2026-05-12). Tools read the DB and the computed JSON artifacts.

lib
@messai/ai-chat
embeddings
HF Inference Router, live
Local dev stack
developer machine
supabase start · :54322/:54323
pnpm dev:zones · 3000/3002/3003/4321

Supabase CLI local stack (Postgres 17 + Studio + Auth + Realtime + Storage) since 2026-05-07. Docker compose for the ML engine on :8001. PDF store on disk is a cache of R2 since 2026-08-14 (pnpm papers:hydrate).

Vercel · 4 projects, one repo
apps/site · messai-site
Vercel zone · marketing
Astro 5 + Vite · 25 pages
/about /learn /solutions /manual …
live · ~8s build

Marketing and docs surface. No DB access, no API routes.

pages
25 .astro (measured 2026-09-03)
dev port
4321
rewrite block
site · 17 entries
apps/web · messai-ai
Vercel zone · product
Next.js 16.3.0-canary.70 · webpack
84 page.tsx · 0 route.ts
live · entry zone

Owns /, /research, /papers, /parameters, /protocols, /admin, /dashboard, /hunter, /insights, /datasets, /experiments, /runs, /survey. The only zone with rewrites. Reads the DB in server components; every write is a fetch to /api.

pages
84 (measured 2026-09-03; docs say 81 @ 2026-07-07)
bundler
webpack — Turbopack wedges on this module graph
build
~5 min
dev port
3000
apps/lab · messai-lab
Vercel zone · lab simulator
Next.js 16.3.0-canary.70 · --webpack
16 page.tsx · R3F + Three.js
live · 3D

/lab/*, /lab/dac (electrochemical DAC workspace, 2026-08-25), /models/*, /predictions/*, /methodology, /parameters/sweep. Owns the heavy 3D tree.

pages
16 (docs said 10 @ 2026-07-06)
dev port
3002
rewrite block
lab · 10 entries
apps/api · messai-api
Vercel zone · API
Next.js 16.3.0-canary.70 · Turbopack
224 route.ts · only DB writer
live · writes

All /api/* handlers. Heavy server deps live here: Sentry, AI SDK, bullmq, ioredis, AWS SDK (R2), pdf-parse. Only zone that writes to Postgres.

routes
224 (measured 2026-09-03; 206 @ 2026-05-20)
caching
force-dynamic default; 12 ISR routes revalidate=60
build
~3.2 min
dev port
3003
newest
admin OpenAlex gap search + GapSearchRun (2026-09-03, feat branch only — not on development yet)
next-auth
auth
GitHub · Google OAuth

NextAuth v4 with the Prisma adapter (two adapters installed — a known P1 cleanup). Sessions, accounts and verification tokens are Prisma models.

Shared code · libs/
@messai/* shared libs · 14
Nx libraries
ui · database · ml · ai-chat
electrochemistry · electrode-3d · mes-3d
component-catalog · dac · polymer-kinetics
platform-health · testing · lab-cad · research-agents

Anything used by more than one zone lives here; apps never import each other's src/. Wired via tsconfig paths, transpilePackages, and workspace:* deps.

  • libs/shared/ui @messai/ui — design system + Tailwind preset, sharp corners only
  • libs/shared/ml @messai/ml — runFullPrediction, per-class routing, priors v2 reader
  • libs/shared/ai-chat — chat tools, live embed via HF router
  • libs/shared/electrochemistry — analytical predictForSystem, 5 physics families
  • libs/shared/dac — electrochemical DAC model (2026-08-25)
  • libs/shared/electrode-3d, mes-3d, component-catalog — 3D recipes + catalog data
  • libs/feature/lab-cad, libs/feature/research-agents
@messai/database
data-access lib
libs/data-access/database
Prisma 6.11 · server-only entry

Single Prisma client. Server contexts import @messai/database/server; there is no db export. Client-safe Zod schemas only outside server modules.

models
117 · 85 enums · 51 migrations
schema
prisma/schema.prisma
Computed JSON artifacts
build-time artifacts
apps/web/public/data/computed/*
priors v1+v2 · calibration · dag-snapshot

Prebuilt outputs of the quality/ML scripts, committed to git and served statically. API routes under research/* and parameters/[slug]/* read DB-first with these as fallback (60s ISR).

  • hierarchical-priors-v2.json — served by /api/ml/predict
  • hierarchical-priors-v1.json — parameter pages, wastewater overlay
  • research/calibration.json (npe-health.json is expected here but absent)
  • protocols/protocol-snapshot.json — 53 steps / 68 edges
Batch operators · no worker host
Operator scripts
client · batch
pnpm tsx scripts/* · python scripts/*.py
dry-run default · --apply --target

All corpus acquisition, extraction, quality refresh and model refits are standalone scripts run by hand or by the weekly Claude routine. There is no worker host.

gate
resolveDbTarget() · scripts/lib/db-target.ts
targets
local | staging | prod (--expect-ref)
Weekly Claude routine
scheduled job · no run evidence in tree
scripts/routines/weekly-ml-audit.sh
6 steps · Monday · commits npe-health.json
batch · weekly

Replaced two GitHub Actions crons (npe-nightly, ml-retrain) that burned ~1,350 min/month. Runs in a Claude remote environment; commits only npe-health.json when it changed. No npe-health.json exists in the repo, so no committed run has landed yet; the 2026-09-03 commit that added step 5b was a dry-run.

  • 1 install Bayes deps (PyMC, sbi, torch)
  • 2 simulator Liu-Logan ±25% gate — hard fail
  • 3 SBC audit N=500
  • 4 compose npe-health.json
  • 5 quality refresh + calibration + GP-SCM fit; 5b two-tier effects refit + OpenAlex gap-search dry-run (added 2026-09-03)
  • 6 platform-health collector
spec
docs/routines/weekly-ml-audit.md
GitHub Actions
CI · manual
workflow_dispatch only · $0 spend
changesets · sync-public-mirrors · fit-priors-v2

No lint/test/build CI. Validation is local via husky pre-commit/pre-push and pnpm pre-deploy. Three workflows remain, all hand-triggered.

  • changesets.yml — version PRs
  • sync-public-mirrors.yml — push open-source/* to Messai-io/MESS-*, publish npm/PyPI via OIDC
  • fit-priors-v2.yml — OOM-safe priors refit
State · storage & queues
Supabase Postgres 17 + pgvector
primary store
prod ref ytgmbbwtawldnunndcnd
staging ref tzyxbgewzgpangsbhoar
live · prod

One schema, three environments (local :54322, staging, prod). Pooled DATABASE_URL :6543 for the app, DIRECT_URL :5432 for migrations and bulk. Staging and prod share the pooler host — writes gate on the project ref.

papers
23,568 local · 23,569 prod (2026-07-06/07)
EPD rows
196,522 · 25,988 modelable (2026-07-07)
vector col
ResearchPaper.abstract_embedding · bge-large 1024d
dead URLs
STAGING_DATABASE_URL, PRODUCTION_DATABASE_URL (db.prisma.io)
Cloudflare R2 · messai-papers
object store
pdfs/ · 7,784 PDFs (2026-08-14)
source of truth for PDF bytes
live · since 2026-08-14

Local papers/pdf-storage/objects/ is only a cache. Hydrate before any job that reads PDFs off disk.

hydrate
scripts/papers/hydrate-pdfs-from-r2.ts
verify
scripts/storage/r2-verify-backups.ts
column
ResearchPaper.r2Key
Upstash Redis + BullMQ
queue · not deployed
apps/api/src/lib/jobs/ · mock processors
dark · dormant

Scaffold only. Processors are TODOs (embeddings write Math.random(), extract_parameters returns 1250). No worker host exists and Vercel cannot run a persistent worker. The 2026-05-30 design retires it in favor of GitHub job-DAG + after().

Inference & models
AI Gateway + model providers
external LLM APIs
Anthropic Haiku 4.5 (default)
Gemini 2.5 Flash · Groq Llama 3.3 70B · Ollama

Chat and extraction go through the Vercel AI Gateway with provider fallback gateway → gemini → groq → ollama (local). Extraction costs ~$0.05–0.10 per paper.

extractor
scripts/extraction/simple_value_extractor.ts
env
AI_GATEWAY_API_KEY, GOOGLE_GENERATIVE_AI_API_KEY, GROQ_API_KEY
HF Inference Router
external inference
BAAI/bge-large-en-v1.5 · 1024d
live query embedding
live · per query

Semantic search embeds the query live via router.huggingface.co, then pgvector cosine on abstract_embedding. Batch backfill uses the Docker TEI service instead.

call site
libs/shared/ai-chat/src/lib/embed.ts
env
HUGGINGFACE_API_KEY
GP-SCM on Fly.io · messai-gp-scm
inference service · deployment not verified
services/ml-engine/Dockerfile.gpscm · :8002
12/27 MFC nodes fitted (2026-07-31)
batch · env-gatedlive · ?hybrid=true

Gaussian-process structural causal model over the named parameter DAG. Reached from /api/ml/predict behind the hybrid opt-in via GP_SCM_SERVICE_URL. A blank value darkened it for ~7 weeks in 2026-07; the client now names that case in its fallback reason. Whether the Fly app is up and the prod env is non-blank was not checked from this machine. Trains with scikit-learn, not torch. Served pickle carries zero non-zero physics means, so no redeploy was warranted after the 2026-07-31 fit.

served pkl
gp-scm-fitted-named-db-2026-06-07.pkl
trainer
services/ml-engine/train_gp_scm.py
ML engine (Docker, local/batch)
batch services
services/ml-engine · :8001
embed · nougat · pgmpy · npe · training/
batch · pnpm ml:dev

FastAPI + Python: batch embeddings (HF TEI), Nougat OCR, pgmpy fitted posteriors (cached JSON, no live server), NPE simulator, hierarchical priors trainers, conformal calibration. Runs from docker-compose.ml.yml with its own postgres/redis/mlflow for experiments.

dockerfiles
Dockerfile · .gpscm · .nougat · .npe
Request and batch paths · 26 edges
FromToPathKind
Researcher browserapps/webHTTPS messai.iolive request path
apps/webapps/siterewrite /about /learn …live request path
apps/webapps/labrewrite /lab /models …live request path
apps/webapps/apirewrite /api/* · all writeslive request path
AI chat & agentsapps/apiPOST /api/chatlive request path
apps/web@messai/databaseserver readsserver-component / build-time read
apps/lab@messai/databaseserver readsserver-component / build-time read
apps/api@messai/databasereads + writeslive request path
@messai/databaseSupabase PostgresPrisma · pooled :6543live request path
apps/apiCloudflare R2PDF byteslive request path
apps/apiUpstash Redis + BullMQenqueue (mock)dormant
apps/apiAI Gatewaychat + toolslive request path
apps/apiHF Inference Routerembed querylive request path
apps/apiGP-SCM on Fly.io?hybrid=truelive request path
apps/apiComputed JSON artifactsJSON fallback · ISR 60sserver-component / build-time read
apps/web · lab · api@messai/* shared libsimports @messai/*server-component / build-time read
apps/apinext-authsessionslive request path
Operator scriptsSupabase Postgressync scripts · DIRECT_URL :5432batch or script
Operator scriptsCloudflare R2hydrate / migratebatch or script
Operator scriptsAI Gatewayextraction LLM callsbatch or script
Weekly Claude routineComputed JSON artifactsrefits → commits JSONbatch or script
Weekly Claude routineGP-SCM on Fly.ioGP-SCM fitbatch or script
GitHub ActionsComputed JSON artifactsfit-priors-v2batch or script
GitHub ActionsML engineworkflow_dispatchbatch or script
ML engineSupabase Postgresbatch embed → pgvectorbatch or script
Computed JSON artifactsapps/webstatic /data/*server-component / build-time read
23,568
papers
local 2026-07-06
7,784
PDFs on R2
2026-08-14
196,522
extracted values
2026-07-07
25,988
modelable
2026-07-07

Four lanes, read top to bottom. Every script is dry-run by default and needs --apply --target to write. The scheduled part is lanes 3 and 4: the weekly Claude routine refits priors, effects, calibration and GP-SCM. Discovery, download and LLM extraction are still hand-run; the GitHub job-DAG that would automate them is a design, not code.

1 · Acquisition — scripts/acquisition/weekly_pipeline.sh (7 steps, dry-run by default)
OpenAlex
external API · discovery
from_updated_date cursor
papers/staging/last-sync.json

The only stage nobody pushes to you, so it is poll-based. Weekly step 2 discovers new DOIs since the cursor; search-gaps-openalex.ts (2026-09-03) runs gap queries into papers/acquisition/gaps-<date>.json and the admin route records a GapSearchRun.

scripts
search_openalex_underrepresented.py · search-gaps-openalex.ts
source tag
ResearchPaper.source = openalex-weekly
Acquire PDFs
scripts · acquisition
paperscraper · curl_cffi
Unpaywall · PMC XML → PDF
batch · manual

Bounded to 2,000 per tier per run. curl_cffi impersonates Chrome to get past MDPI/Wiley/ACS/Elsevier 403s. 2,075 no-DOI papers remain unreachable by this path.

scripts
download_paperscraper.py · download_curl_cffi.py · download_unpaywall.py · xml_to_pmc_pdf.py
Promote to canonical store
script
promote_to_canonical.py
papers/pdf-storage/objects/ (cache)
batch · manual

Content-addressed by pdfHash. Since 2026-08-14 the bytes live on R2 and the local store is a cache; migrate-local-to-r2.ts / backfill-r2-keys.ts keep ResearchPaper.r2Key populated.

R2 · messai-papers
object store
7,784 PDFs · 2026-08-14

Source of truth for PDF bytes.

Sync papers → DB
script · DB write
sync-papers-to-db.ts
pdfHash link · taxonomy · gated
batch · gated

Creates or links ResearchPaper rows. Mutation gate mirrors the whole pipeline: dry-run unless --apply --target {local|staging|prod}; prod requires --expect-ref.

ResearchPaper
table
23,568 rows (local 2026-07-06)
82 relation edges in schema

Hub of the schema. Carries pdfStoragePath, r2Key, aiSummary, taxonomy columns (primarySystemType TEXT), demo content, and the pgvector abstract_embedding.

2 · Parse & extract — scripts/extraction, scripts/overnight
Hydrate PDFs
script
pnpm papers:hydrate
batch · pre-job

Pulls the needed PDFs from R2 into the local cache before any extraction cohort runs.

Parse
OCR / parse
Nougat OCR (Docker) → .mmd
PMC XML → text (pmc-xml-to-text.ts)
batch · manual

Nougat runs as an overnight batch via pnpm nougat-batch. 1,444 BioC-PMC XML files are parsed by an existing parser but not yet re-extracted (cost-gated ~$70–150).

server
services/ml-engine/nougat_server.py
out
papers/nougat/<doi>.mmd
v2 value extractor
LLM extraction
simple_value_extractor.ts → values_v2
Haiku → gemini → groq → ollama
batch · $0.05–0.10/paper

The active extractor (v1.6 orchestrator deprecated 2026-05-08). Writes JSON under papers/extracted/<aa>/<sha>/; every run must end with its sync exiting 0 (Rule §7). 9,511 papers done as of the last full run.

fallback
docs/extraction/provider-fallback-2026-05-09.md
smoke
scripts/extraction/phase6-smoke-test.ts
Overnight jobs A–E
orchestrator
pnpm overnight:run
XML→v2 · equations · figures · overview
batch · manual

Five sequential jobs against the local corpus, ending in job E which syncs all four artifact types to the DB. Ollama-backed with provider fallback; manifest survives SIGTERM.

runbook
docs/runbooks/overnight-pipeline.md
Legacy Claude/Gemini scripts
older extractors
extract-tables-claude.py
digitize_figures_gemini.py · …
dark · ad hoc

Still present and runnable; hard-code claude-sonnet-4-6 (P1 cleanup). Populate ExtractedTableData / ExtractedFigureData / ExperimentalContext.

Sync extractions → DB
script · DB write
sync-values-v2-to-db.ts · sync-*-to-db.ts
with extractorVersion · snippet · confidence
batch · gated

Lands rows with full provenance. audit-disk-vs-db.ts fails pre-deploy on disk-only drift. Launch-blocker fields (derivationMethod, uncertainty±, catholyteBuffer) are proper columns, not JSONB.

ExtractedParameterData
table
196,522 rows · 25,988 modelable
80.7% NULL canonical_slug

Per-paper extracted values with conditions, canonical slug, SI-normalized numericValue, verifierPassed. The binding constraint on modelability is canonicalization coverage, not extraction.

Extracted{Table,Figure,Paper}Data
tables
ResearchPaperEquation · ConditionSet
ExperimentalContext · ReactorGeometry

Typed sidecar tables linked by paperId. ConditionSet carries phAnolyte/phCatholyte/reference electrodes as columns; ReactorGeometry has ~47 typed columns.

3 · Quality, embeddings & model refits — scripts/quality/refresh-all.sh (11 steps) + services/ml-engine/training
refresh-all.sh
quality pipeline
1 ConditionSet · 2 canonical slugs · 3 SI normalize
4 physics · 5 outliers · 5b verifier · 10 disk-vs-DB
batch · after each batch

Idempotent, dry-run by default. Volumetric densities route to their own kinds so they never pool with areal values in priors. Propagation order is local → staging → prod, additive and ref-checked.

normalizer
scripts/quality/normalize-to-si.py
canonicalize
canonicalize-name.ts
Priors & calibration refit
ML training
hierarchical_priors_v2.py (served) + v1 (pages)
split-conformal calibration · PPC on held-out
batch · weekly routine

Steps 6–9 of refresh-all. v2 is Student-t, per-class stratified, nutpie sampler; the full fit must use the OOM-safe batched runner. Never ship a FAST-mode artifact. Both v1 and v2 must be refit together.

batched
training/run_priors_v2_batched.py
benchmark
bes-benchmark-v1 (paper-disjoint, 2026-09-02)
Within-paper effects + KG
derived research
two-tier effect fit, DB-fed (2026-09-03)
build-correlations-db.ts · LearnedDagEdge
batch · feat branchbatch · routine 5b

Effects now fit from the DB, not just the table CSV, with gap records feeding the OpenAlex gap search. KG correlation matrix is slug-keyed Pearson + Spearman.

embed_papers.py
embedding backfill
bge-large-en-v1.5 → abstract_embedding
HF TEI Docker :8001
batch · manual

Batch backfill of the pgvector column. If the service is down the embedding is silently skipped and nothing retries.

GP-SCM fit
ML training
train_gp_scm.py → .pkl → Fly image
batch · weekly

12/27 MFC nodes fitted after the present-parent fix; binding constraint is joint-data coverage.

4 · Serving
Computed artifacts (git)
static JSON
apps/web/public/data/computed/*
priors v1/v2 · calibration · npe-health · dag-snapshot

Committed to main by the routine only on substantive change; a commit deploys all four Vercel projects.

apps/api routes
serving
/api/ml/predict · /api/research/* · /api/parameters/*
DB-first, JSON fallback, ISR 60s on 12 routes
live · 224 routes

runFullPrediction reads priors v2, conformal calibration, OOD detection and runtime physics validation. Non-MFC classes route to the analytical predictor; MFC keeps the data-tuned path.

UI · apps/web · apps/lab
surfaces
/research · /parameters · /protocols · /lab
live · browser

Semantic search: live query embedding → pgvector cosine → hydrate rows. New extractions appear instantly; prebuilt artifacts need a redeploy or pnpm regenerate.

corpus-refresh.yml (proposed, not landed)
proposed automation
GitHub job-DAG · cron 0 2 * * 1
discover→acquire→extract→sync→embed
dark · design 2026-05-30

The design in docs/corpus-refresh-architecture.md. Not implemented: the workflows directory holds only changesets, sync-public-mirrors and fit-priors-v2. Acquisition and extraction remain manual; only refits are scheduled (weekly Claude routine).

Pipeline edges · 26 edges
FromToPathKind
OpenAlexAcquire PDFsnew DOIsbatch or script
Acquire PDFsPromote to canonical storedownloaded/batch or script
Promote to canonical storeR2upload · r2Keybatch or script
R2Sync papers → DBmanifestbatch or script
Sync papers → DBResearchPaperinsert / linkbatch or script
R2Hydrate PDFspull cachebatch or script
Hydrate PDFsParsePDFsbatch or script
Parsev2 value extractor · Overnight jobs.mmd / textbatch or script
ParseLegacy Claude/Gemini scriptsad hocdormant
v2 value extractorSync extractions → DBvalues_v2.jsonbatch or script
Overnight jobs A–ESync extractions → DBjob Ebatch or script
Legacy scriptsExtracted{Table,Figure,Paper}Datatables · figures · contextdormant
Sync extractions → DBExtractedParameterData · sidecar tablesrows + provenancebatch or script
ExtractedParameterDatarefresh-all.shreads EPD · writes slug, SI value, verifierPassedbatch or script
refresh-all.shPriors & calibration refitsteps 6–9batch or script
Priors & calibration refitWithin-paper effects + KGthenbatch or script
Within-paper effects + KGGP-SCM fitthenbatch or script
ResearchPaperembed_papers.pyabstractsbatch or script
Priors & calibration refitComputed artifacts (git)priors · calibrationbatch or script
Within-paper effects + KGComputed artifacts (git)effects · dagbatch or script
GP-SCM fitapps/api routespkl → Flybatch or script
Computed artifacts (git)apps/api routesJSON fallbackserver-component / build-time read
ExtractedParameterDataapps/api routesDB-first readslive request path
apps/api routesUI · apps/web · apps/labfetch /api/*live request path
embed_papers.pyUI · apps/web · apps/labpgvector cosinelive request path
Within-paper effects + KGOpenAlexgap records → gap search (5b)batch or script
9
packages in open-source/
workspace:*
835
parameters · 15 categories
mess-parameters v0.3.0
349
refs to mess-parameters
apps/libs/scripts
$0
Actions spend
dispatch-only

Since the 2026-04-25 consolidation every package lives in this repo; the public Messai-io repos are read-only mirrors pushed by hand at release time. mess-parameters is the heavyweight: the product reads its ontology at build time and the DB syncs extracted values back into it. Dataset blobs never enter git or npm: the catalog ships manifests, the bytes live on Hugging Face Hub.

open-source/ · 9 packages · workspace:*

PackageVersion · license · filesWhat it isMirror · refs
mess-parametersv0.3.0 · CC-BY-4.0 · 1,444
Standardized parameter ontology and analysis tools: 835 parameters across 15 categories. data/parameter-definitions-rich.json (pinned v0.2.0 for fixtures) plus SCIENTIFIC_INTEGRITY.md (power density CoV ≈1,285%).
  • parameters/*.md ontology
  • data/paper-parameter-values.csv — mirror of DB extractions
  • scripts/papers/ Snakefile pipeline (honors PAPERS_ROOT)
@messai-io/mess-parameters
github.com/Messai-io/MESS-Parameters
349 refs in main repo
mess-materialsv0.2.0 · CC-BY-4.0 · 104
DFT-computed material properties for electrodes, membranes, catalysts with Materials Project provenance, Pourbaix stability, elasticity.
  • seeded by scripts/sync-materials-to-db.ts, seed-materials-bootstrap.ts
@messai-io/mess-materials
github.com/Messai-io/MESS-Materials
69 refs in main repo
mess-microbesv0.1.0 · MIT · 41
Curated microorganisms relevant to MES: electrogens, electrotrophs, community partners and intentional negative controls. Catalog = 28 microbes (corrected 2026-06-30).
  • export-microbes-to-json.ts · audit-microbe-field-coverage.ts
@messai-io/mess-microbes
github.com/Messai-io/MESS-Microbes
34 refs in main repo
mess-datasets-catalogv0.1.0 · CC-BY-4.0 · 612
Catalog and classifications for open MES datasets (Zenodo, Figshare), slug-keyed to parameters and materials. Metadata only; blobs on HF Hub.
  • 23 records · 108 files · 27 MB on HF (2026-04-25)
  • per-record manifest.json with download_url + checksum
@messai-io/mess-datasets-catalog
github.com/Messai-io/MESS-datasets
6 refs in main repo
mess-simulationsv0.1.0 · MIT · 31
Physics-based simulation and modeling tools for MES.
@messai-io/mess-simulations
github.com/Messai-io/MESS-Simulations
4 refs in main repo
mess-agentsv0.1.0 · MIT · 27
Multi-agent research orchestration framework.
@messai-io/mess-agents
github.com/Messai-io/MESS-Agents
1 refs in main repo
mess-hypothesesv0.1.0 · MIT · 28
Research-gap identification and hypothesis generation. No chat tool wraps it yet (P2 gap).
@messai-io/mess-hypotheses
github.com/Messai-io/MESS-Hypotheses
1 refs in main repo
mess-learningv0.1.0 · CC-BY-4.0 · 22
Educational content and calculators.
@messai-io/mess-learning
github.com/Messai-io/MESS-Learning
1 refs in main repo
mess-methodspython · pyproject · 32
Python package of MES methods (tests/, src/). Published to PyPI by the mirror workflow.
PyPI mess-methods
github.com/Messai-io/MESS-Methods
1 refs in main repo
messai-ai monorepo · github.com/samfrons/messai.ai · open-source/ is the source of truth
apps/* + libs/*
consumers
4 zones · 14 shared libs
generate-parameter-fixtures.ts → /data/parameters/<slug>.json

The product reads packages at build time (fixtures), at seed time (materials, microbes), and at runtime (datasets manifests → HF Hub).

Supabase Postgres
store
ParameterDefinition · Material · Microbe
ExtractedParameterData

ParameterDefinition is edited in the DB and synced back to mess-parameters; materials and microbes are seeded from their packages.

Releases · pnpm changeset
release mechanism
.github/workflows/changesets.yml
tools/lint-package-shape.ts
batch · dispatch

Changesets version PRs; canonical package shape enforced by the linter. Rollback = git revert in open-source/* then re-run the mirror.

sync-public-mirrors.yml
GitHub Action
workflow_dispatch · package input
pnpm publish + PyPI via OIDC
batch · manual, $0

Pushes open-source/<pkg> to its Messai-io repo with a tag, then publishes. Kept on GitHub because OIDC trusted publishing only works there. Auto-triggers were stripped to hold Actions spend at $0.

Public · downstream, read-only
github.com/Messai-io/MESS-*
public GitHub org
9 public mirror repos
read-only downstream

MESS-Parameters, MESS-Materials, MESS-datasets, MESS-Agents, MESS-Hypotheses, MESS-Learning, MESS-Microbes, MESS-Methods, MESS-Simulations. Contributions from the public flow back via docs/contributing-from-public.md.

npm · @messai-io/*
registry
8 packages · pnpm publish
PyPI · mess-methods
registry
OIDC trusted publishing
Hugging Face Hub
blob host
datasets/messai-io/mess-datasets
108 files · 27 MB · 23 records (2026-04-25)

Large dataset blobs (PDFs, CSVs, images) referenced by download_url in each manifest; consumers verify the md5 checksum after download.

External researchers
audience
install packages · open issues · PRs on mirrors
Sync, publish and consume paths · 13 edges
FromToPathKind
mess-parametersapps/* + libs/*fixtures at buildserver-component / build-time read
Supabase Postgresmess-parametersdefs + values syncbatch or script
mess-materialsSupabase Postgresseed Materialserver-component / build-time read
mess-microbesSupabase Postgresseed Microbeserver-component / build-time read
mess-datasets-catalogapps/* + libs/*/datasets manifestsserver-component / build-time read
Releases · pnpm changesetsync-public-mirrors.ymlversion PR mergedbatch or script
sync-public-mirrors.ymlgithub.com/Messai-io/MESS-*git push + tagbatch or script
sync-public-mirrors.ymlnpm · @messai-io/*pnpm publishbatch or script
sync-public-mirrors.ymlPyPI · mess-methodsmess-methods onlybatch or script
mess-datasets-catalogHugging Face Hubblobs · download_urlbatch or script
apps/* + libs/*Hugging Face Hubfetch + checksumlive request path
github.com/Messai-io/MESS-*External researchersclone / issuesserver-component / build-time read
npm · @messai-io/*External researchersinstallserver-component / build-time read
117
models
116 on development
85
enums
prisma/schema.prisma
146
related model pairs
from the schema relation list
51
migrations
2025-08-07 → 2026-09

ResearchPaper (82 relation fields) is the hub; User (38), Experiment (34), Microbe (26), ParameterDefinition and ConditionSet (22 each) follow. Only ResearchPaper carries a pgvector column. Schema rule §8: any field that a UI filter, ML feature or WHERE clause touches is a typed column, never JSONB.

DomainModelsNames
Papers & corpus12ResearchPaper · PaperRelationship · PaperMergeAudit · ResearchCluster · ResearchTrend · AnomalousPaper · ResearchPaperEquation · PaperParameter · PaperMaterial · PaperMicrobe · MetagenomicPaperLink · GapSearchRun
Extraction14ExtractedParameterData · ExtractedTableData · ExtractedFigureData · ExtractedPaperData · ExtractionJob · ExtractionProvenance · ExperimentalContext · ConditionSet · ParameterConditionLink · ParameterObservation · ParameterProvenance · ReactorGeometry · OperatingCondition · MFCDesign
Parameters & knowledge19ParameterDefinition · ParameterEdge · ParameterPrior · PriorStratumAxis · ParameterTemplate · ParameterClassification · CustomField · ElectrochemicalParameter · BiologicalParameter · EnvironmentalParameter · OperationalParameter · MaterialParameter · KnowledgeNode · KnowledgeEdge · LearnedDagEdge · SymbolicLaw · SubstrateClassification · DataClassificationRule · BufferChemistryProfile
Materials7Material · MaterialClassification · MaterialMicrobeAffinity · MaterialPaperCrossref · ElectrodeFormulation · AiMaterialDefinition · AiGeometryRecipe
Microbes & biofilm15Microbe · MicrobeCytochrome · MicrobeShuttle · MicrobePerformanceMeasurement · MicrobeEngineeringEvent · MicrobeOmicsStudy · MicrobeReference · MicrobeSubstrateProductPair · MicrobeKineticConstant · EETPathwayAssignment · MetagenomicAnalysis · AiMicrobeDefinition · BiofilmSample · BiofilmCreepCurve · MicroelectrodeProfile
Experiments & lab21Experiment · ExperimentCollaborator · ExperimentEvent · ExperimentPaper · Run · Measurement · LabConfiguration · LabConfigurationSnapshot · LabWastewaterQueryLog · MethodologyPreset · Workflow · WorkstreamArtifact · SimulationReplay · SimulationResult · Model · KineticCurve · PolarizationPoint · EISPoint · ElectrochemicalKinetic · MESProduct · MESProductYield
ML & agents9Prediction · CrossSystemPrediction · CalibrationResult · TrainedModelLineage · EvalGateRun · EvalResult · AgentRun · Hypothesis · IntegrationMapping
Datasets5Dataset · DatasetMeasurement · DatasetCurvePoint · DatasetParameter · DatasetMaterial
Users & admin11User · Account · Session · VerificationToken · Permission · Team · Project · BetaSignup · SurveyResponse · SharedView · AuditLog
Feedback4Feedback · FeedbackAnalytics · FeedbackNotification · WastewaterFeedback

What the diagrams draw dotted or red. Severity follows the platform docs; states were re-checked against CLAUDE.local.md and the working tree on 2026-09-03. Resolved items (chat bifurcation, force-dynamic audit, priors JSON validity) are listed separately.

SevGapCurrent stateWhat unblocks it
P0Extraction & acquisition are not scheduledweekly_pipeline.sh, simple_value_extractor.ts run by hand; corpus-refresh.yml is a design (2026-05-30) with no code. Only refits run weekly (Claude routine).Land the GitHub job-DAG or a scheduled routine step for discover → acquire → extract with the budget cap.
P0BullMQ layer is a mockapps/api/src/lib/jobs processors are TODOs; no worker host; Vercel cannot host a persistent worker.Archive the scaffold; use after() / Vercel Queues for the real-time ingest plane.
P1Canonicalization coverage, not extraction, bounds modelability80.7% of EPD rows have NULL canonical_slug (84,243 numeric rows recoverable). Modelable 25,988 vs ceiling ≈40–50k.Extend alias maps / canonicalize-name.ts; re-run refresh-all steps 2–3.
P1Prod FK under-population~33k ExtractedParameterData rows in prod lack parameterDefinitionId (affects FK/DAG read paths, not the modelable count).Ref-gated backfill, staging first.
P1Papers ↔ Materials / Microbes orphanedanodeMaterials / cathodeMaterials are strings; PaperMaterial / PaperMicrobe junctions mostly empty.Backfill junctions from extraction; add /api/materials/[id]/papers.
P1Unread corpus1,444 BioC-PMC XML files parsed but not re-extracted (~$70–150); 2,075 no-DOI papers unreachable; MDPI/Wiley/ACS/Elsevier 403s.Cost-gated v2 run over XML; curl_cffi path for no-DOI URLs.
P1Confidence and integrity caveats not surfacedconfidence columns exist; UI shows point estimates. SCIENTIFIC_INTEGRITY.md (power-density CoV ≈1,285%) has no UI banner.Add confidence to /api/parameters/* responses; collapsible callout on /parameters/*.
P1GP-SCM fits are data-limited12/27 MFC nodes fitted; new fits are interpolation artifacts (energy_efficiency flat at 43.02%). toc_removal has 1 joint observation.Targeted paid re-extraction for biofilm/biomass nodes; do not lower --min-samples.
P2No feedback / retraining loopPrediction table exists; no /lab widget records {predicted, actual}; no drift-triggered refit.Feedback widget → POST /api/predictions; weekly regenerate if drift > 5%.
P2Chat tools missing for 4 packagesquery-methods, query-hypotheses, query-datasets, query-learning not wrapped.One tool per package in @messai/ai-chat.
P1GP-SCM can be silently dark in prodA blank GP_SCM_SERVICE_URL on Vercel disabled the pillar for ~7 weeks in 2026-07 while it looked wired (libs/shared/ml/src/gp-scm-client.ts). Whether prod has a non-blank value today was not verified from this machine.Check the Vercel env on messai-api; the client now reports the blank-value case as its fallback reason.
P1Weekly routine output never landedNo npe-health.json exists anywhere in the tree, so either the routine has never run with --commit or its commit step has not fired. The routine script itself is present and gained step 5b on 2026-09-03.Run bash scripts/routines/weekly-ml-audit.sh --commit once by hand; confirm the scheduled session exists in Claude Code on the web.
P2Dev hygieneTwo NextAuth Prisma adapters installed; react-query in devDependencies. (Legacy extractors already read ANTHROPIC_MODEL with a claude-sonnet-4-6 default — that item from the 2026-05 docs is resolved.)Pick @auth/prisma-adapter; move dependency.
Resolved since the docs were written (verify before re-flagging)
  • Chat bifurcation — one canonical /api/chat (2026-05-12).
  • Priors JSON validity — v1 and v2 parse cleanly (commit 715b63e96).
  • Per-class ML routing — 6 non-MFC classes route to the analytical predictor (commits 4778f8589, bef7acedd); lab UI still passes proxy inputs.
  • Harmonization quick wins merged (PR #501): +434 modelable; propagated to staging and prod 2026-07-06/07.
  • GP-SCM all-parents dropna — present-parent fit measured 9 → 12 nodes (2026-07-31, branch fix/gpscm-present-parent-fit, unmerged).
Sources: docs/platform/00-architecture.md (verified 2026-07-06) · 05-data-flow.md (2026-05-06) · 06-infrastructure-gaps.md · docs/corpus-refresh-architecture.md (2026-05-30) · docs/routines/weekly-ml-audit.md · CLAUDE.md · CLAUDE.local.md · prisma/schema.prisma. Route, model and lib counts measured from the working tree on 2026-09-03; corpus counts carry their own dates — re-measure before quoting (scripts/check-actual-database-stats.ts). The 146 related-pair count is derived from the artifact's embedded schema relation list.