What this page is. We commissioned an external methodological review of our cost model and belief elicitation form. The reviewer identified several real structural problems. We are sharing the critique openly — and our responses — because we think transparency about limitations is as important as the model itself. (This page reviews the cost model and elicitation design; the subject-matter evaluation of the underlying forecast is Evaluation 2.)
Context. The review was conducted in early 2026 based on the model and form as they existed at that time. Both have since been updated — some issues below are already addressed; others are work in progress. The model will continue to improve after the May 8 workshop.
The full report is available on GitHub: CM_model_deep-research-report.md →
The reviewer's bottom line: "The project is already good enough to support productive discussion, but not yet good enough to serve as the authoritative modeling-and-elicitation backbone for coordinated expert convergence." We agree with this assessment. The workshop is designed for productive discussion and deliberation — not to produce a single authoritative posterior. The limitations below matter, and we are naming them rather than papering over them.
What the model and process already do well
- Public and transparent — code, formulas, and parameter choices are all exposed and annotatable
- Uncertainty-aware — Monte Carlo simulation with explicit distributions rather than point estimates
- Linked to deliberation — designed to elicit beliefs before and after evidence review, not just poll after
- Explicitly caveated — documentation notes limitations including the static snapshot, missing geography, and ad hoc dependence structure
- Discussion-connected — Hypothes.is annotation, GitHub Discussions, and participant forms all feed back into the model and site
Key structural issues in the model
▸ Target quantity: "edible kg" vs "wet-weight cell mass at harvest" Addressed
The model page defines the output as "pure cultured chicken cell biomass (wet weight, at harvest)" — factory-gate, before texturisation or blending. The beliefs form asks for "average production cost per edible kg of cultured chicken meat." These are actually the same accounting object (harvested wet cell mass = the edible component before any downstream processing), but the wording difference created potential for semantic confusion.
▸ Missing supplemental recombinant proteins (albumin, transferrin, insulin) Structural gap
The model's "growth factors" parameter covers FGF-2, IGF, TGF-β and similar signaling proteins. But three other recombinant proteins used in cell culture media — albumin, transferrin, and insulin — are not yet separately modeled. According to GFI's 2023 analysis, these together are expected to account for the vast majority of recombinant protein production volume in CM, and under optimistic (efficient media) conditions could contribute ~$1/kg. Under pessimistic conditions — high media use, no substitution — the contribution could be $5–50/kg or more, potentially larger than several items the model currently tracks.
Albumin is the most important: it's used at gram-per-liter concentrations, and animal-free alternatives (rice-derived, yeast-derived) are expensive today but improving. Transferrin is lower volume but expensive. Insulin is lowest volume and recombinant insulin is relatively affordable.
▸ Dependence structure: the single "maturity" factor Structural gap
The model uses one latent "maturity" variable to correlate technology adoption rates, financing costs (WACC), and reactor choice. This prevents absurd combinations — e.g., technology succeeding while financing remains stuck — but asks one synthetic dimension to stand in for sector maturity, supply-chain development, regulatory learning, and financing all at once. The reviewer notes this can create hidden coupling and double-counting.
Better practice would separate at least three latent factors: technical bioprocess maturity, supply-chain / input maturity, and financing / regulatory maturity — with user-selectable correlation among them. This would also make the dependence structure itself an object of scrutiny rather than a baked-in assumption.
▸ Sensitivity analysis: dollar-swing ranking, not variance decomposition Addressed in source
The model's tornado chart shows how much the mean cost changes when each parameter moves from its 10th to its 90th percentile — a conditional-mean dollar-swing statistic. The reviewer notes this is not a Sobol variance decomposition: bars for correlated parameters can overlap and should not be summed. An earlier version of the simplified-view text incorrectly said these parameters "contribute less than 10% of the variance," which was the wrong framing for a swing statistic.
A proper global sensitivity analysis (Sobol first-order and total-effect indices) would answer a fundamentally different question: what fraction of the output variance is explained by each input, accounting for nonlinear and interaction effects? This is methodologically superior but requires dedicated re-simulation infrastructure.
Elicitation design issues
▸ Form depth: lightweight survey, not full distribution elicitation Elicitation
The focal question asks for a median and optional 80% credible interval. Most technical subquestions ask for point estimates. The Sheffield Elicitation Framework and similar structured protocols recommend eliciting full probability distributions for each quantity — using fixed-interval (chips-and-bins) or variable-interval (quartiles) methods — and include calibration training, piloting, and revision rounds.
This form is much simpler: a structured survey without calibration training, not a full elicitation instrument. That's a real limitation for the goal of updating model priors rigorously.
▸ Pooling: mixed expert types need separate reporting Elicitation
The workshop mixes TEA authors, industry operators, bioprocess researchers, evaluators, and animal-welfare stakeholders. Good-practice elicitation recommends reporting subgroup beliefs separately before attempting any pooled synthesis — subgroup disagreement is information, not noise. If TEA specialists and industry operators systematically disagree, that disagreement is exactly what the workshop is designed to surface.
Performance-weighted aggregation (e.g., Cooke's Classical Model) goes further: using calibration seed questions with known answers to weight experts by accuracy and informativeness. This requires preparing calibration questions in advance — something we did not do for this workshop.
▸ No calibration training Elicitation
Formal expert elicitation protocols typically include training on cognitive biases in probability estimation — anchoring, overconfidence, availability heuristics — along with practice questions and feedback. This helps experts produce better-calibrated probabilities. We are not doing this at this workshop: participants are not being paid, time is limited, and there is no trained facilitator running a formal protocol.
Some of the workshop participants (e.g., David Manheim) have forecasting/elicitation backgrounds that partially substitute for formal training. But the group overall has not been calibrated.
What the workshop itself is designed to do
The workshop is not primarily an aggregation exercise — it's a deliberation exercise. The design (pre-workshop beliefs → evidence review → live discussion → post-workshop belief update) is exactly the right structure for collective learning, even if the elicitation instrument is simpler than a formal protocol. The value of having TEA authors, industry operators, and evaluators in the same room for structured disagreement is not diminished by the form being a lightweight survey.
What the workshop cannot do: generate a calibrated posterior that can stand on its own as an authoritative quantitative estimate of CM_01. For that, we would need the full formal apparatus. We hope participants understand this going in.
Interested in going further?
If there is enough interest from participants — and if we secure resources — we would design a follow-up exercise with:
- Calibration training — practice questions with known answers, feedback on bias
- Full distribution elicitation — chips-and-bins or quartile method for each key parameter
- Individual-before-group — beliefs elicited privately before any group interaction
- Subgroup comparison — separate posteriors for TEA specialists, industry operators, evaluators
- Compensation — paid time for serious engagement
- Facilitator — trained in structured expert elicitation (e.g., SHELF protocol)
If you're interested in participating in this, or in helping lead it, please say so in the beliefs form's "About You" section — and feel free to email contact@unjournal.org.