Regression Test Suite
Full item-level results from the 600-item regression test suite. The suite checks calculation behaviour and, where a comparable baseline is available, rating stability and CO2 reproducibility across three test cases. These are internal checks, separate from the CAB’s independently drawn live sample.
Run of 2026-10-01 — suite v2.7.0 on the dev-validation stack, EDB 2026.1 (edb-to-bafu v1.12.0, released 2026-09-30). This is the state prepared for review; no earlier state is offered for comparison, and the first comparison on this page will be 2026.1 as reviewed against 2026.2, once that release exists. Suite v2.7 contains 600 declarations drawn against the released edb-to-bafu scope and carries all 965 reachable base products (4 of the 969 are unreachable by design). 599 of 600 items computed; the other one returns no value as expected, and there are no unexpected failures — caterer 249/250, label/retail 200/200, full specification 150/150. The one is a caterer recipe carrying a branded product that this release withholds, so it has no LCA inventory here; it stays in the suite for coverage. The run is pinned to the committed suite code and to the inventory and method files of edb-to-bafu v1.12.0. It replaces the run of 2026-09-30 on edb-to-bafu v1.11.1. Against that run the change is material (EACS-12 §12.4.2, 593 items computed on both: 9.4 % rating changes, 13.2 % CO2 changes over 10 %): items with a notable peanut share rise strongly, several curry and coconut dishes rise, tomato sauces fall by 20–42 %, and the remaining items move by a median of −0.7 %. Six recipes that had no inventory on v1.11.1 (pea-protein meat substitutes and a curry sauce) now compute. v2.7 answers the review of v2.6: 243 caterer recipes and all 150 full-specification products were re-issued under new ids, each naming its v2.6 predecessor. Caterer recipes were re-read so that each reads as a coherent dish, full-specification products got plausible undeclared amounts, origins and names, and base products that no item carried were added where they fit. Representativeness: the 250 caterer recipes are drawn from real food-service production data (74 customer namespaces, 2024 onwards, deduplicated on composition) and stratified by namespace, √-weighted with an 8% cap per stratum — the case spans the recipe population Eaternity receives rather than reproducing its frequencies, and no single customer exceeds 7.2% of it. Every ingredient string is replaced by a different real term and amounts are jittered, so no item carries its source’s strings or amounts. The 200 label and 150 full-specification items descend from disjoint draws on a frame of openly published CH/DE/AT product declarations, stratified to a Swiss convenience-retail assortment and reconstructed with fantasy brand names and synthetic GTINs.
Local Oracle ↔ EOS Parity
The offline Local Oracle reproduces the production EOS climate score across the 600-item suite.
Vintage: This parity measurement predates suite v2.4 and was made on the v1.2.0 population, not the 600 items shown above; it is evidence about the oracle, not about this run’s recipes. Parity was measured against the full production run of 2026-08-28 (a separate production reference, pre-re-baselining; not the 2 September validation-environment run) with the oracle's EOS-pick cache rebuilt from that same run. Known production defects at this vintage, stated rather than hidden: poultry by-products (chicken shank, chicken edible offal, paste broth) resolve to a ~22 g CO₂/kg, 0 kcal target — a matching defect (eos#21) whose fix will be the first post-baseline change record; and production changed its data state on 26.8. (EDB package 2026-05-04 → 2026-06-08 plus ~954 new name→term matchings), so the 2026-08-25 run measured the May package (butter 11.75, milk 1.95 kg/kg) and this run the June package (9.0 / 1.43, matching the July anchor; eos#22 tracks whether a cache recompute contributed) — which is why the validation baseline was re-pinned to this run on 2026-08-28. Ratings are compared across two ladders — the oracle rates on the declared 2026 benchmark ladder, production still rates on the pre-2026 anchor — so the rating-match figure reads as a declared-vs-live gap, not as oracle error. The re-pin against the re-baselined stack on the validation environment follows the final run.
Median absolute divergence by case
Largest divergences
| Recipe | Case | EOS (g) | Oracle (g) | Δ |
|---|
These are the worst-case items in the resolved sample — the median across all 598 resolved items is far smaller. Rows where EOS reports a small whole-gram value (e.g. 1–3 g per portion) show a large percentage for a sub-gram absolute difference; the genuine disagreements are the high-gram rows where a single ingredient resolved to a different background than EOS picked.
Items not used for the median
| Item | Case | Reason | Why excluded |
|---|
| ID ▲ | Title ▲ | Case ▲ | CO2 (g) ▲ | Rating ▲ | Computed ▲ | Deviation ▲ |
|---|