process_log initializing...
A practitioner journal · biweekly · free

We document the real work of building process digital twins.

Every issue covers one domain, one problem: the parameters that actually drive outcomes, what the model told us, and what we had to rethink. Written by the engineers who build MycoTwin, OliveTwin, and CargoTwin.

// no email required — every issue is published in full right here


archive All Issues
ISS.04 The Batch You Never Have to Run Cross-Twin ISS.03 Mass Balance First: Why Yield Models Fail Without It OliveTwin ISS.02 Why Mushroom Yields Are Harder to Predict Than You Think Mushroom ISS.01 What a Process Digital Twin Actually Is (For People Who Build Real Things) Methodology

Issue 04 · Methodology · Cross-Twin

The Batch You Never Have to Run

Totec Labs·8 min read

Here's a number worth sitting with: a bad batch costs you the batch. Not the process, not the recipe — the physical inputs. The olives already pressed, the milk already in the vat, the substrate already inoculated. You don't get those back by learning from the mistake. You get a lesson and an invoice.

A bad simulation costs you nothing. Maybe a few seconds of compute, maybe a fraction of a cent if there's an API call involved. That asymmetry is the entire economic case for a digital twin, and it's easy to lose sight of it once you're deep in the modeling work and focused on prediction accuracy instead. Accuracy matters. But the reason accuracy matters is that it lets you fail somewhere that doesn't hurt.

Monitoring tells you. A twin lets you rehearse.

We made this distinction in Issue 01 — a dashboard tells you what happened, a twin tells you what will happen if you change something. What we didn't spell out then is that "tells you what will happen" is really shorthand for something more useful: it lets you run the version of the batch that goes wrong, on purpose, before you ever run the version that's real.

That's rehearsal, not prediction. And once you build a twin with rehearsal in mind instead of just forecasting in mind, you start designing different features.

OliveTwin's scenario presets are a direct example. Alongside the live batch monitor, there's a set of pre-built conditions you can load and predict against without touching a single real olive: Optimal, High Temp, High O₂, Conservative, Late Season. That High Temp / High O₂ scenario isn't decoration — it's the specific failure mode a mill operator is most likely to hit by accident (malaxation running hot, headspace not purged) modeled ahead of time, so the first time they see what it does to predicted polyphenol retention is on a screen, not in a batch of oil that's already pressed.

MycoTwin does the same thing with its contamination risk assessment. You can run that check before you inoculate anything — before spawn goes into substrate, before a facility commits blocks, labor, and weeks of climate-controlled space to a run. Contamination in mushroom cultivation isn't a "reduce yield a little" failure mode. It's often a "lose the whole block, sometimes lose the room" failure mode. Catching an elevated-risk condition in the model costs you a look at a number. Catching it three weeks into spawn run costs you the block.

Calibration is rehearsal's foundation

Rehearsal only works if the thing you're rehearsing against tells the truth about your specific process — not a textbook average of everyone's process. This is where DairySim's "Golden Batch" calibration earns its keep. Instead of running on generic kinetic constants, it tunes the hidden rate parameters — the fermentation kinetics governing pH drop, curd firmness development, κ-casein cleavage — against your facility's own historical batch data, minimizing the error between what the model would have predicted and what your vats actually did.

From the build
A model running on textbook kinetics will get you in the right neighborhood. A model calibrated to your specific starter culture, your specific vat geometry, your specific ambient conditions will tell you what your process actually does — not what cheddar generally does.

Picture a cheesemaker running a new order through DairySim before committing the vat — testing rennet dosage against set time the way they normally would by feel, except this time watching the model show curd firmness crossing the target threshold four minutes earlier than the standard schedule calls for, because the calibration has already picked up that this facility's vat runs slightly warmer than the reference conditions the generic recipe assumes. Cut on the textbook schedule and you're under-set. Cut on the calibrated prediction and you're not. The difference between those two outcomes was available before the rennet ever went in — but only because the model had been calibrated against real history first.

What "solving it digitally" actually means

None of this is about the model being smarter than the practitioner. Every one of these twins is built on parameters an experienced operator already understands — malaxation temperature, contamination risk, cut time. The twin's job isn't to know more than you. It's to let you test the version of the decision you're not sure about, as many times as you need to, in a space where being wrong doesn't cost anything.

That's the actual value of the rehearsal scenarios, the risk checks, the calibration step. Not "the model will tell you what to do." More precisely: the model gives you a place to be wrong first, cheaply, repeatedly, until the answer you're about to act on in the real world is one you already trust.

If your twin only ever runs on the batch you're actually about to commit to, you're leaving most of its value on the table. Build in the ability to run the batch you're worried about — the hot one, the contaminated one, the underdosed one — before you build the one you're planning on. The cost of finding out you were wrong should never be higher on a screen than it is in the tank, the room, or the mill. If it is, the twin isn't doing its job yet.

← Back to all issues
Issue 03 · Methodology · OliveTwin

Mass Balance First: Why Yield Models Fail Without It

Totec Labs·7 min read

Ask ten process engineers what a yield model does and nine of them will describe the same thing: feed it your process parameters, it spits out a percentage. That's not wrong. It's just incomplete in the specific way that eventually breaks the model.

A yield number on its own doesn't tell you where the rest of the mass went. And "the rest of the mass" is where every yield model that later falls apart on you actually failed first — you just didn't notice, because the model never asked the question.

The question a percentage can't answer

We ran into this building OliveTwin. Early versions modeled polyphenol yield as a single output: olives in, phenol content out, done. It worked fine right up until a mill operator changed one thing — swapped a 3-phase centrifuge for a 2-phase — and the model's predictions quietly stopped matching reality. Same olives, same malaxation profile, same everything on paper. Different number every time.

The model wasn't wrong about the chemistry. It was blind to where the mass was going.

The core distinction
A yield percentage is an output. Mass balance is the constraint that makes the output trustworthy.

Polyphenols in an olive batch don't just end up in the oil or "not in the oil." They split three ways — into the oil itself, into the vegetation water (wastewater), and into the pomace — and how they split depends on equipment you might not think of as a process variable. A 2-phase centrifuge adds less water to the must than a 3-phase system, which changes the dilution the polyphenols are sitting in during separation, which changes how much ends up retained in the oil versus washed out into the wastewater stream. Same phenol chemistry. Different partition. Different final number.

If your model only tracks "yield," this is invisible. If it tracks the mass balance — where every unit of phenol content actually goes — it's not just visible, it's diagnosable.

What actually has to balance

Once you're tracking partition instead of a single output, you also have to account for the fact that polyphenols degrade during processing, not just redistribute. That degradation isn't random — it's kinetic, and it's exactly the kind of thing a mass balance forces you to model honestly instead of fudging into a fixed "loss factor."

degradation rate (k, 20°C ref)s⁻¹
activation energy (Ea)kJ/mol
O₂ reaction order (n)dimensionless
yield rate constant (k_y)min⁻¹

OliveTwin models degradation as a temperature-driven rate — a reference rate at 20°C, scaled by an activation energy term the way any Arrhenius-governed reaction would be, plus a separate reaction order on oxygen exposure. That second part matters more than people expect: it's why N₂ blanking during malaxation is a real lever and not a gimmick. Cut the oxygen headspace, and you're not adjusting a vague "freshness" setting — you're directly reducing the rate constant on a specific degradation pathway that's competing with extraction for the same phenol mass.

This is also why the alpha-weighting control in OliveTwin — the slider between "maximize polyphenol" and "maximize yield" — actually means something instead of being a dial you eyeball. It only works as a real trade-off because the model is tracking the same conserved phenol mass splitting between two outcomes: staying concentrated in a smaller volume of higher-quality oil, or spreading across a larger volume of oil at lower phenol density. You can't optimize a trade-off you haven't modeled as a balance. You can only guess at it.

Same principle, different process

This isn't an olive oil quirk. It's the same requirement showing up in every twin we've built.

MycoTwin's biological efficiency metric — fresh mushroom yield relative to dry substrate input — is a mass balance ratio by definition. If a grower's actual BE consistently undershoots the model, that's not "the model is off," that's a signal something's leaving the system unaccounted for: incomplete colonization, moisture loss during incubation, contamination eating substrate the model assumed was going to fruiting bodies. The ratio doesn't just predict yield. It tells you where to go looking when the prediction is wrong.

DairySim's yield coefficient — grams of curd solids produced per gram of lactose consumed — does the same job for a fermentation-driven process. It's the same accounting discipline applied to a biological conversion instead of a physical partition, and it fails the same way if you skip it: you get a number that's right on average and useless for diagnosis the one time it's wrong.

Before you add parameter sensitivity, before you add an AI advisor layer, before you add anything that makes a model feel smarter — check that it balances. Every unit of mass or concentration you start with has to be accounted for somewhere at the end: in the product, in a byproduct stream, in a documented loss mechanism. If your model can't tell you where something went when a prediction misses, it's not a twin. It's a curve fit wearing a twin's clothes, and curve fits break the moment your process conditions drift outside the range they were fit on.

Get the balance right first. Everything else you build on top of it — sensitivity, calibration, recommendation engines — inherits that foundation's honesty or inherits its blind spots.

← Back to all issues
Issue 02 · Mushroom · MycoTwin

Why Mushroom Yields Are Harder to Predict Than You Think

Totec Labs·8 min read

Ask an experienced grower what drives yield and you'll get a confident answer: humidity, temperature, fresh air exchange. All true. What's missing is that these variables don't act independently — they interact, and the interactions are where most yield loss actually happens.

We built MycoTwin because we kept seeing the same pattern across grow operations: an operator adjusts one variable in response to a problem, the problem seems to resolve, and three weeks later a different problem shows up that traces back to the same change. The feedback loop between cause and effect in mushroom cultivation is long enough that intuition alone struggles to close it.

The variables that actually move the needle

In our modeling work, four parameters accounted for the overwhelming majority of yield variance across the cultivation cycle: CO₂ concentration during pinning, relative humidity during fruiting, fresh air exchange rate, and substrate moisture content at spawn. Individually, growers manage these reasonably well. The problem is the interaction terms.

CO₂ (pinning)800–1,200 ppm
RH (fruiting)85–95%
FAE rate4–6 exchanges/hr
substrate moisture60–65%

Here's the part that surprises most operators: pushing CO₂ down to encourage pinning while humidity is already at the low end of range doesn't just risk two separate problems — it compounds into a single failure mode where pins initiate but abort before maturing. Neither variable alone predicts this. The interaction does.

From the model
In our simulation runs, batches with CO₂ and RH both drifting toward their low-range boundaries simultaneously showed yield loss nearly triple what either variable predicted independently. This is the kind of failure mode that's invisible until you model the interaction directly.

Why this matters for scaling

A single grow tent forgives a lot — you're physically present, you notice the environment shifting, you correct in real time. Scale to multiple rooms or a commercial facility, and that hands-on correction disappears. Variables you used to catch by feel now need to be modeled, because you can't be in four rooms at once.

This is the actual case for simulation in a niche like mushroom cultivation. It's not about replacing grower intuition — it's about giving that intuition something to check itself against before a bad batch, not after.

If you're scaling past a single grow room, the interactions between your environmental variables matter more than any single setpoint.

← Back to all issues
Issue 01 · Methodology

What a Process Digital Twin Actually Is (For People Who Build Real Things)

Totec Labs·6 min read

"Digital twin" gets used two very different ways, and the gap between them causes most of the confusion practitioners run into when they hear the term.

In enterprise contexts — aerospace, automotive, large-scale manufacturing — a digital twin usually means a live, sensor-fed replica of a physical asset, continuously synced to real-world data. Think a jet engine's twin updating in real time from thousands of embedded sensors. That's a legitimate and valuable thing. It's also completely out of reach for a 500-block mushroom operation or a mid-size olive mill, and frankly, not what most small operators need.

The definition that actually matters here

For the processes we work on, a digital twin means something more specific and more useful: a predictive model of a process, built from real domain data, that lets you test a change before you commit resources to it. No live sensor feed required. No IoT infrastructure. Just a validated model you can run scenarios against.

The core distinction
A dashboard tells you what already happened. A digital twin tells you what will happen if you change something — before you change it for real.

This distinction matters because it changes what you should expect to pay for, and what you should expect the tool to do. A monitoring dashboard tracks your CO₂ levels. A digital twin tells you what happens to your yield if you change your CO₂ setpoint by 200ppm, before you touch anything.

What makes a twin trustworthy

Three things, in our experience: the model has to be built on real domain data, not generic assumptions. The assumptions baked into the model need to be documented, not hidden. And the tool needs to be honest about its limitations — what it can predict confidently, and where the real world will diverge from the simulation.

The underlying principle stays the same across every domain we've built for: understand the process first, validate the model against real data second, ship the tool third.

That's the standard we hold every twin we build to — and it's the lens this publication comes from.

← Back to all issues