Regression Test Suite
Full item-level results from the 600-item regression test suite. Each run validates rating stability and CO2 reproducibility across three certification cases.
Local Oracle ↔ EOS Parity
The offline Local Oracle reproduces the production EOS climate score across the 600-item suite.
Median absolute divergence by case
Largest divergences
| Recipe | Case | EOS (g) | Oracle (g) | Δ |
|---|
These are the worst-case items in the resolved sample — the median across all 595 resolved items is far smaller. Rows where EOS reports a small whole-gram value (e.g. 1–3 g per portion) show a large percentage for a sub-gram absolute difference; the genuine disagreements are the high-gram rows where a single ingredient resolved to a different background than EOS picked.
Items not used for the median
| Item | Case | Reason | Why excluded |
|---|
| ID ▲ | Title ▲ | Case ▲ | CO2 (g) ▲ | Rating ▲ | Calculation ▲ | Deviation ▲ |
|---|
Tributary Review Priority — Recipe-Grounded
Which LCI tributaries actually contribute the CO2 of the 600 recipes we evaluate — each ingredient weighted by its real recipe CO2, its base product recovered via the oracle's own sediment resolver and exact-joined to a tributary. This is our rooting in reality: where review changes the answers that matter.
| # | Tributary | Recipe CO2 share | Instances | Priority |
|---|
Tributary Review Priority — Catalog Change
The complementary full-catalog view: which tributaries the EDB→BAFU migration changed the most across all 2,331 base products — ranked by share of the CO2 difference (|BAFU − ecoinvent|). This catches products not yet in any recipe; the recipe-grounded view above is the reality check on what's actually consumed.
| # | Tributary | Products | CO2 change share | Magnitude | Mean conf. | Priority |
|---|