Cultured Chicken Cost Model · CM Workshop (May 8, 2026)

Model Limits & Critique

An external methodological review — what the model gets right, what it doesn't, and what we're doing about it.

What this page is. We commissioned an external methodological review of our cost model and belief elicitation form. The reviewer identified several real structural problems. We are sharing the critique openly — and our responses — because we think transparency about limitations is as important as the model itself. (This page reviews the cost model and elicitation design; the subject-matter evaluation of the underlying forecast is Evaluation 2.)

Context. The review was conducted in early 2026 based on the model and form as they existed at that time. Both have since been updated — some issues below are already addressed; others are work in progress. The model will continue to improve after the May 8 workshop.

The full report is available on GitHub: CM_model_deep-research-report.md →

The reviewer's bottom line: "The project is already good enough to support productive discussion, but not yet good enough to serve as the authoritative modeling-and-elicitation backbone for coordinated expert convergence." We agree with this assessment. The workshop is designed for productive discussion and deliberation — not to produce a single authoritative posterior. The limitations below matter, and we are naming them rather than papering over them.

What the model and process already do well

Key structural issues in the model

▸ Target quantity: "edible kg" vs "wet-weight cell mass at harvest" Addressed

The model page defines the output as "pure cultured chicken cell biomass (wet weight, at harvest)" — factory-gate, before texturisation or blending. The beliefs form asks for "average production cost per edible kg of cultured chicken meat." These are actually the same accounting object (harvested wet cell mass = the edible component before any downstream processing), but the wording difference created potential for semantic confusion.

What we've done: Added an explicit accounting boundary note to both the beliefs form and the About page CM_01 box: "The 'edible kg' refers to the harvested cell biomass itself (wet weight), factory-gate, before texturisation, blending, or packaging."
▸ Missing supplemental recombinant proteins (albumin, transferrin, insulin) Structural gap

The model's "growth factors" parameter covers FGF-2, IGF, TGF-β and similar signaling proteins. But three other recombinant proteins used in cell culture media — albumin, transferrin, and insulin — are not yet separately modeled. According to GFI's 2023 analysis, these together are expected to account for the vast majority of recombinant protein production volume in CM, and under optimistic (efficient media) conditions could contribute ~$1/kg. Under pessimistic conditions — high media use, no substitution — the contribution could be $5–50/kg or more, potentially larger than several items the model currently tracks.

Albumin is the most important: it's used at gram-per-liter concentrations, and animal-free alternatives (rice-derived, yeast-derived) are expensive today but improving. Transferrin is lower volume but expensive. Insulin is lowest volume and recombinant insulin is relatively affordable.

What we're doing: The Expert Distribution section of the beliefs form (E4) now asks participants to estimate the total albumin + transferrin + insulin cost contribution per kg of cell biomass. These elicited estimates will be used to update the model's priors and add a "supplemental proteins" cost component. The beliefs form's Expert Mode already captures this; the main model will be updated to include it as a separate line item.
▸ Dependence structure: the single "maturity" factor Structural gap

The model uses one latent "maturity" variable to correlate technology adoption rates, financing costs (WACC), and reactor choice. This prevents absurd combinations — e.g., technology succeeding while financing remains stuck — but asks one synthetic dimension to stand in for sector maturity, supply-chain development, regulatory learning, and financing all at once. The reviewer notes this can create hidden coupling and double-counting.

Better practice would separate at least three latent factors: technical bioprocess maturity, supply-chain / input maturity, and financing / regulatory maturity — with user-selectable correlation among them. This would also make the dependence structure itself an object of scrutiny rather than a baked-in assumption.

Near-term plan: Multi-factor dependence is on the model development roadmap. As a first step, the model could offer three modes: independent, sector-coupled (current), and strongly coupled — letting users see whether headline results are robust to the correlation assumption. We will implement this in the model repository after the workshop.
▸ Sensitivity analysis: dollar-swing ranking, not variance decomposition Addressed in source

The model's tornado chart shows how much the mean cost changes when each parameter moves from its 10th to its 90th percentile — a conditional-mean dollar-swing statistic. The reviewer notes this is not a Sobol variance decomposition: bars for correlated parameters can overlap and should not be summed. An earlier version of the simplified-view text incorrectly said these parameters "contribute less than 10% of the variance," which was the wrong framing for a swing statistic.

A proper global sensitivity analysis (Sobol first-order and total-effect indices) would answer a fundamentally different question: what fraction of the output variance is explained by each input, accounting for nonlinear and interaction effects? This is methodologically superior but requires dedicated re-simulation infrastructure.

What we've done: The source text now accurately describes the tornado chart as a "conditional-mean swing statistic, not a variance decomposition." The chart is labeled accordingly. In September 2026 we added a sensitivity methods page that works through why the bars are hard to interpret (coupled inputs, arbitrary tails, tail-driven means, learning versus setting an input) and runs alternatives on the same scenario: adjustable tails and statistics with standard errors, conditional profiles, expected uncertainty after learning one input, rank regression, one-at-a-time interventions, and an exploratory Shapley-effect estimate. A Sobol decomposition on independent primitive inputs and value-of-information analysis remain open.

Elicitation design issues

▸ Form depth: lightweight survey, not full distribution elicitation Elicitation

The focal question asks for a median and optional 80% credible interval. Most technical subquestions ask for point estimates. The Sheffield Elicitation Framework and similar structured protocols recommend eliciting full probability distributions for each quantity — using fixed-interval (chips-and-bins) or variable-interval (quartiles) methods — and include calibration training, piloting, and revision rounds.

This form is much simpler: a structured survey without calibration training, not a full elicitation instrument. That's a real limitation for the goal of updating model priors rigorously.

What we're doing: The Expert Distribution Mode section of the form (below the main questions) asks for p10/median/p90 distributions for five key model parameters: basal media $/L, growth factor price $/g, growth factor dosage g/kg, and supplemental protein cost $/kg. This is richer than a point estimate. For those willing to go further, Metaculus supports full distributions. We're also asking in the form whether participants would be interested in a more rigorous future process.
▸ Pooling: mixed expert types need separate reporting Elicitation

The workshop mixes TEA authors, industry operators, bioprocess researchers, evaluators, and animal-welfare stakeholders. Good-practice elicitation recommends reporting subgroup beliefs separately before attempting any pooled synthesis — subgroup disagreement is information, not noise. If TEA specialists and industry operators systematically disagree, that disagreement is exactly what the workshop is designed to surface.

Performance-weighted aggregation (e.g., Cooke's Classical Model) goes further: using calibration seed questions with known answers to weight experts by accuracy and informativeness. This requires preparing calibration questions in advance — something we did not do for this workshop.

What we're doing: The "About You" section now has nine expertise categories (TEA author, bioprocess researcher, industry operator, industry analyst, AW researcher, AW funder, evaluator/forecaster, general interest, other) designed for meaningful subgroup comparisons. We commit to reporting subgroup distributions — not just a pooled posterior — in the workshop synthesis. Performance-weighted aggregation is not feasible for this round but is listed as a future extension.
▸ No calibration training Elicitation

Formal expert elicitation protocols typically include training on cognitive biases in probability estimation — anchoring, overconfidence, availability heuristics — along with practice questions and feedback. This helps experts produce better-calibrated probabilities. We are not doing this at this workshop: participants are not being paid, time is limited, and there is no trained facilitator running a formal protocol.

Some of the workshop participants (e.g., David Manheim) have forecasting/elicitation backgrounds that partially substitute for formal training. But the group overall has not been calibrated.

What we're doing: The beliefs form now asks whether participants have previously done formal expert elicitation with calibration training, and whether they'd be interested in a more rigorous process in future. If there's sufficient interest and funding, we would design a follow-up with proper calibration training, compensation, and a trained facilitator. The workshop itself — with deliberation before and after evidence review — partially mitigates anchoring through group discussion.

What the workshop itself is designed to do

The workshop is not primarily an aggregation exercise — it's a deliberation exercise. The design (pre-workshop beliefs → evidence review → live discussion → post-workshop belief update) is exactly the right structure for collective learning, even if the elicitation instrument is simpler than a formal protocol. The value of having TEA authors, industry operators, and evaluators in the same room for structured disagreement is not diminished by the form being a lightweight survey.

What the workshop cannot do: generate a calibrated posterior that can stand on its own as an authoritative quantitative estimate of CM_01. For that, we would need the full formal apparatus. We hope participants understand this going in.

Interested in going further?

If there is enough interest from participants — and if we secure resources — we would design a follow-up exercise with:

If you're interested in participating in this, or in helping lead it, please say so in the beliefs form's "About You" section — and feel free to email contact@unjournal.org.

Note on the model site: The cost model is maintained in a separate repository and continues to be updated independently. Changes to model structure — adding supplemental proteins, multi-factor dependence, global sensitivity analysis — will be implemented there. This page covers the workshop-facing beliefs elicitation aspects.