Senior Staff Software Engineer · Tech Lead

Mohit
Nagpal

I build systems that turn ambitious growth ideas into dependable, measurable platforms—at massive scale.

Portrait of Mohit Nagpal
Uber Growth TechEngineering leadership
14years in software development
2technology companies: Uber & Microsoft
∞focus on systems built to scale

Engineering at the intersection of scale, growth, and people.

I’m a Senior Staff Software Engineer and Tech Lead at Uber Growth Tech, with 14 years of software development experience across Uber and Microsoft.

My work spans the technical and organizational layers required to make complex platforms succeed: reliable distributed systems, high-throughput data pipelines, rigorous experimentation, practical GenAI adoption, and engineering leadership that helps teams move with clarity.

CurrentUber Growth Tech
Career experienceUber · Microsoft

Complex systems, made useful.

01

Massive-scale distributed systems

Architecture & reliability

02

Data pipelines

Quality & throughput

03

GenAI adoption

From idea to practice

04

Experimentation

Evidence-led growth

05

Engineering leadership

Direction & leverage

Independent inquiry, grounded in engineering.

Alongside my industry work, I’m an independent researcher. I’m interested in the ideas, evidence, and tools that can improve how modern software systems are built and evaluated.

GitHub work@MohitNagpaldce ↗
ORCID0009-0006-1732-058X
02 Evaluating LLM Outputs in Production: A Practical Framework Grow golden datasets from production failures, correct LLM-as-judge biases, and gate merges on per-slice eval scores. Read articleClose article

Shipping an LLM feature is easy. Knowing whether it still works six months, three model versions, and two prompt refactors later is the hard part. Over the last few years of building LLM-backed systems that serve real traffic, I've converged on a framework that actually catches regressions before users do — without requiring a research lab's budget.

Start with a golden dataset, not vibes

Every eval practice starts with the same artifact: a set of input-output pairs that encode "this is what good looks like." The mistake I see most often is treating the golden dataset as a one-time artifact assembled by an engineer in an afternoon. A useful golden set is grown, not written.

The highest-signal examples come from production incidents and user corrections. Every time someone flags a bad response, fixes an answer, or rephrases a query after getting a bad result, that interaction is a candidate for the dataset — with the corrected output as the expected answer. This keeps your evals anchored to the failures your users actually experience rather than the failures you imagine.

Aim for breadth over depth early: 100–200 examples covering your distinct task types (extraction, summarization, classification, grounded QA, open-ended generation) beats 1,000 near-duplicate QA pairs. Tag each example with its task type and difficulty, because aggregate pass rates lie — a model can hold 92% overall while collapsing on one task type, and you won't notice unless you slice.

One more thing: version your golden set. When you add an example because a prompt change fixed a real incident, tag it with the date and the incident. Six months later, when someone asks "why does this prompt look weird," the answer is in the dataset history.

LLM-as-judge works — if you know what it's lying about

Human grading doesn't scale, so most teams graduate to LLM-as-judge: a stronger model scores outputs against a rubric. This works surprisingly well for structured tasks — our judge-based scores correlated with human ratings at roughly 0.8 on extraction and summarization tasks. But the judge has systematic biases, and if you don't correct for them, your evals will quietly certify bad behavior.

Self-preference bias. A judge model favors outputs that resemble its own generation style. If your judge is from the same model family as your generator, expect inflated scores. Mitigate it: use a different model family for judging than for generation, and periodically calibrate against human labels on a held-out slice. If judge-human agreement drifts below your threshold (we use 0.75 correlation as the floor), the judge needs recalibration, not the pipeline.

Verbosity bias. Judges reward longer answers, even when the rubric says "be concise." This is one of the most replicated findings in eval research, and it bites hardest on summarization evals — the judge will prefer the 8-sentence summary over the correct 3-sentence one. Counter it with explicit rubric instructions ("penalize answers that include information not requested"), and add a length-normalized check: flag any output more than 2x the reference length for human review.

Position bias. When the judge compares two outputs (A vs. B), it favors whichever came first — or second, depending on the model. The fix is trivial and widely skipped: run every pairwise comparison in both orders and only count wins that are consistent across both. Inconsistent orderings are ties. Yes, this doubles judge cost. It's still cheaper than shipping a "better" prompt that wasn't actually better.

Also worth internalizing: judges are worst at exactly the tasks where you most want them — open-ended generation with subjective quality bars. Use LLM judges for structured, rubric-checkable tasks, and reserve humans for the fuzzy ones.

Human-in-the-loop: sample like you mean it

You can't have humans review everything, so sampling strategy is the whole game. Random sampling wastes human attention on easy cases; the defects live in the tails.

Uncertainty sampling works well in practice: have the judge emit a confidence score alongside each grade, and route the lowest-confidence grades to humans. In our experience, the bottom 10% of judge-confidence outputs contained the majority of real defects — a 10x leverage on reviewer time.

Stratified sampling guards against blind spots: sample a fixed minimum per task type, per language, per user cohort — even when volumes are low. The bug that only affects 0.3% of traffic can still be the one that ends up in a support ticket from your largest customer.

And always keep a canary slice: 20–50 production queries per day that humans review no matter what. This is your ground truth against which you validate the judge itself. If judge-human agreement on the canary slice drops, you have a judge problem, and every eval result since the last healthy check is suspect.

Regression-test your prompts like code

A prompt change is a code change. Treat it like one: every pull request that touches a prompt, a model version, or a retrieval config should run the eval suite and show the diff in scores — per task type, not just aggregate.

This catches a class of failure I call capability whack-a-mole: you fix the model refusing to answer financial questions, and it starts refusing to answer legal questions. Aggregate accuracy stays flat; per-slice accuracy reveals the trade. Without per-slice regression gates, prompt engineering degrades into moving bugs around.

Pin your model versions in the eval harness. Model providers update weights under the same version string more often than they'd like you to know. Record the actual model identifier (and, when available, the system fingerprint) in every eval run so a mysterious score shift is attributable. When a pinned model degrades without any change on your side, that's a provider-side change — and it's worth knowing before your users tell you.

Set thresholds as gates, not decorations. A suite that "runs" but never blocks a merge is theater. We gate on two conditions: no task-type slice may drop more than 2 absolute points, and the canary agreement floor must hold. Everything else is informational.

Track evals in CI, and make the history visible

Evals that only run locally get forgotten. Wire them into CI so every prompt or model change produces a score diff, and store the results in a time-series store — not a spreadsheet, not a Slack thread. You want to be able to answer "when did summarization quality start declining?" with a chart, not an archaeology expedition.

The metrics worth trending: per-task-type pass rates, judge-human agreement on the canary slice, output length distributions (a sudden shift often precedes a quality shift), and cost per eval run (judge calls are real money; a suite that costs $40 per run gets run less often than one that costs $4).

A few operational numbers from experience: a solid starting suite is 150–300 golden examples with judge grading at roughly $0.01–0.05 per example per run — cheap enough to run on every prompt PR. Human review budget goes entirely to the canary slice plus uncertainty-sampled outliers, typically a few dozen items per week for a mid-size feature. That ratio — broad automated coverage, narrow deep human coverage — is the whole framework in one sentence.

The bottom line

Eval quality compounds. The teams that do this well aren't the ones with the fanciest judge prompts; they're the ones whose golden datasets grow from production incidents, whose judges are calibrated against humans on an ongoing basis, and whose prompt changes face the same regression discipline as code changes. Start with a golden set built from real failures, add a calibrated judge with known biases corrected, sample humans where the judge is weakest, and gate your merges on per-slice scores. Everything else is optimization.

03 RAG at Scale: What Breaks When Retrieval Meets Production Traffic Chunking failure modes, hybrid retrieval with reranking, index freshness vs. rebuild cost, embedding drift, and latency/cost budgets. Read articleClose article

RAG demos are magical. RAG in production is a distributed systems problem wearing a trench coat. The failure modes that matter don't show up at 100 queries a day — they show up at 10,000 queries a minute, with a corpus that changes hourly and a latency budget measured in hundreds of milliseconds. Here's what breaks, and what actually fixes it.

Chunking: the decision that quietly determines everything

Chunking strategy is the highest-leverage decision in a RAG system and the least reversible one. Every downstream component — embeddings, index, reranker, the prompt itself — is built on top of chunk boundaries. Get it wrong and you're retuning everything.

The failure modes of naive chunking are specific and predictable. Fixed-size character chunking splits sentences and tables mid-thought; the embedding of a half-sentence is a noisy vector that matches queries it shouldn't. Overly large chunks dilute the embedding — a 2,000-token chunk about five topics retrieves for queries about any of them, and then the generator has to find the relevant paragraph inside a haystack you handed it. Overly small chunks lose the context needed to answer; the retriever finds the right fact but the generator can't use it because the chunk lacks the antecedent three sentences.

What works in practice: chunk on semantic boundaries (sections, paragraphs, table rows) rather than token counts, and store chunk metadata — document title, section heading, position — so the generator can cite and the retriever can filter. A pattern I've seen work well is hierarchical: index small chunks for retrieval precision, but pass the parent section to the generator for context. It costs one extra lookup and fixes the context-loss problem almost entirely.

The non-obvious trap: chunking interacts with your embedding model's context window. If your chunks are longer than what the embedding model was trained to encode well (often well under its advertised max), quality degrades silently — embeddings still come back, they're just worse. Test retrieval quality at your actual chunk sizes, not the model's spec sheet.

Hybrid search: dense retrieval is not enough

Pure dense (vector) retrieval fails on exact-match queries — product codes, error strings, proper nouns, version numbers. A vector index will happily return a semantically similar but factually wrong document for ERR_CONNECTION_RESET because "connection" and "reset" are semantically near lots of things. BM25 (sparse/keyword) search handles these perfectly and fails on paraphrase. You need both.

The standard production architecture is retrieve-then-rerank: a hybrid first stage (dense + BM25, often fused with reciprocal rank fusion) pulls 50–100 candidates cheaply, then a cross-encoder reranker scores each query-document pair properly and picks the top 5–10 for the prompt. The reranker is the single biggest quality lever after chunking — it routinely fixes 10–20% of top-k errors — but it's also the latency bottleneck, since cross-encoders score pairs sequentially.

Two practical notes. First, tune the fusion weights on your own query log, not on a public benchmark. The dense-vs-sparse balance that wins on MS MARCO is not the balance that wins on your users' actual queries. Sample 500 real queries, label relevance for the top candidates, and grid-search the weights. Second, add metadata filters before the vector search, not after: filtering by tenant, document type, or date range at the index level shrinks the candidate space and is nearly free, while post-filtering wastes your top-k slots on documents you were going to discard.

Index freshness vs. rebuild cost

Your corpus changes; your index lags. The question is how much lag you can afford, and the answer depends on the document type. A knowledge base that updates daily can tolerate a nightly rebuild. A system indexing support tickets or incident reports may need minutes.

Full rebuilds are simple and correct but expensive — re-embedding a million documents is a real GPU bill and takes hours. Incremental updates are cheaper but introduce their own problems: deleted documents that linger as ghost vectors, updated documents that exist in two versions with different embeddings, and index fragmentation that degrades recall over time.

The pattern that holds up: incremental updates with periodic compaction. Stream inserts, updates, and deletes into the index continuously, and run a full rebuild on a schedule (weekly or monthly) to clear fragmentation and drift. Keep a tombstone list for deletes so ghost vectors are filtered at query time between rebuilds. And measure recall on a fixed probe set after every rebuild — if incremental indexing is silently degrading, the probe set catches it before users do.

One more freshness trap: embedding model upgrades. Switching embedding models means re-embedding the entire corpus — there is no in-place upgrade, because vectors from different models live in different spaces. This is a migration, not a config change: dual-index during the transition, A/B the retrieval quality, and only cut over when the new index wins on your probe set.

Embedding drift: the silent quality killer

Even without model changes, retrieval quality drifts. The corpus distribution shifts (new product lines, new terminology), user query patterns shift (seasonality, new features), and the alignment between the two erodes. This is embedding drift, and it's silent because nothing errors — recall just decays a few points per quarter until someone notices answers getting worse.

You can't fix what you don't measure, so build a retrieval probe set: 100–200 representative queries with known-relevant documents, run nightly. Track recall@k and MRR. Set an alert on sustained degradation, not single-day noise — retrieval metrics are noisy, and paging someone over a 2-point daily wobble trains everyone to ignore the alert.

When drift is confirmed, the fixes in order of cost: refresh the probe set (maybe the world changed and the probes are stale), tune hybrid weights, add synonyms and query rewriting for new terminology, and only then consider re-embedding or fine-tuning the embedding model. Most drift I've seen was fixed by the first three.

Latency budgets and the cost that compounds

RAG latency is a sum, and every component spends from the same budget: query embedding (~20–50ms), hybrid retrieval (~30–100ms), reranking (~100–300ms for 50 candidates), prompt construction, and generation (the long pole, often 1–3s). If your p99 budget is 2 seconds, the reranker and the generator are fighting over the same slack.

The highest-ROI latency wins: cache embeddings for repeated queries (support bots see heavy query repetition), cap rerank candidates at what the quality curve justifies (the gain from 50→100 candidates is usually marginal), stream the generation so time-to-first-token drops even when total time doesn't, and consider a smaller reranker — a distilled cross-encoder at 100 candidates often beats a large one at 30 within the same latency.

Cost per query compounds the same way latency does. A typical RAG query costs roughly: embedding ($0.0001), vector DB reads (fractional at scale, but provisioned throughput isn't free), reranker inference ($0.001–0.01 depending on model and candidate count), and generation ($0.002–0.05 depending on model and context length). At 10k queries a day, the difference between a lean and a sloppy pipeline is hundreds to thousands of dollars a month — and the dominant term is almost always context length in the generation call. Every unnecessary retrieved chunk you stuff into the prompt is tokens you're paying for on every query. Tight top-k selection isn't just a quality decision; it's the cost decision.

The bottom line

RAG at scale fails at the seams: chunk boundaries that split meaning, dense-only retrieval that misses exact matches, indexes that lag the corpus, embeddings that drift from the query distribution, and latency and cost budgets spent by components that don't know about each other. The throughline is measurement — a probe set for retrieval quality, per-component latency breakdowns, per-query cost tracking — because every one of these failure modes is silent until you instrument it. Build the observability first; the tuning decisions become obvious once you can see what's happening.

04 A/B Testing Done Right: The Statistical Pitfalls Every Engineer Should Know Peeking, sample ratio mismatch, novelty effects, power analysis, guardrail metrics, and pre-registered stop rules. Read articleClose article

Most A/B tests I've seen in industry are decided before the statistics are. Someone peeks at the dashboard on day three, sees a promising lift, and ships it. The test was theater; the decision was vibes. The statistical pitfalls below aren't academic niceties — they're the specific ways teams fool themselves with data, and each one has a concrete fix.

Peeking: the false-positive machine

Here's the mechanism. A standard A/B test is designed for exactly one look at the data, at a pre-committed sample size, with a 5% significance level. That 5% is a contract: if there's no real effect, you'll falsely declare a winner 5% of the time. But if you check the results daily for two weeks and stop the moment p < 0.05, you've run fourteen tests, not one — and your real false-positive rate is closer to 20–30%, not 5%.

This is the single most common way teams ship "wins" that don't exist. The dashboard makes it worse: real-time p-values invite peeking the way a slot machine invites another pull. I've watched teams ship three consecutive "winning" variants, each reverting the last, because every decision was noise.

The fixes, in order of practicality:

  1. Pre-commit to a runtime. Decide the sample size (or duration) before launch, and don't look at the primary metric until it's done. This is the simplest correct approach and the one most teams should use.
  2. Sequential testing methods (group sequential designs, always-valid p-values) if you genuinely need early stopping. These adjust the significance threshold at each look to preserve the overall error rate. They require actual statistical machinery — don't hand-roll them; use a library.
  3. Separate monitoring from decision-making. It's fine to watch for bugs, outages, and SRM (below) during the test. Just don't make ship decisions off interim primary-metric reads.

A useful team norm: the person who can see the live dashboard should not be the person who decides when to stop the test.

SRM: sample ratio mismatch

SRM means your traffic split isn't what you configured — you asked for 50/50 and got 52/48, or the ratio drifts over time. A small mismatch sounds harmless. It isn't: SRM is often the fingerprint of a broken randomization or a filtering bug that biases which users enter the test, and a biased sample invalidates everything downstream.

Common causes I've seen: the experiment assignment happens after a redirect that drops a cookie, so one variant systematically loses mobile users; a bot filter applied to one arm but not the other; a caching layer that serves the control page to users who should be in treatment. The mechanism varies; the signature is the same — the observed split deviates from the expected split beyond what chance explains.

Detection is a chi-square goodness-of-fit test on the unit counts per arm, run continuously. Every experimentation platform worth using has this built in — if yours doesn't, that's a gap to close. When SRM fires, stop and investigate; don't "adjust" for it. There is no statistical correction for a broken randomization. Find the assignment bug, fix it, and restart the test. Shipping a result from an SRM-flagged test is building on a cracked foundation.

Novelty and primacy effects

Users react to change itself, not just to your change. A redesigned checkout flow gets a burst of engagement because it's novel — users click around the new thing — and that burst decays over a week or two. Conversely, a change that disrupts habit (moved buttons, renamed features) gets an initial penalty — the primacy effect — that also decays as users relearn.

The failure mode: you run a test for five days, measure during the novelty window, and ship a "lift" that evaporates by week three. Or you kill a genuinely good change because week-one numbers looked bad during the relearning dip.

The fix is boring and effective: run tests long enough to outlast the novelty window — typically at least one to two full weekly cycles, longer for habit-heavy surfaces. And segment your analysis by exposure cohort: compare users in their first three days of exposure against users with 10+ days. If the effect exists only in the fresh cohort, it's novelty, not a win. If the effect grows with exposure, you may have cut a good test short.

Power analysis and minimum detectable effect

Before launching, answer two questions: how small an effect do you care about (the minimum detectable effect, MDE), and how long until you can reliably see it (the sample size for 80% power at your significance level)?

Teams routinely skip this and then run tests that were doomed from the start — a test powered to detect a 5% lift cannot tell you anything about a 1% lift, and most product changes move metrics by 1% or less. The result is a graveyard of "inconclusive" tests that burned three weeks each, or worse, underpowered tests where a lucky noise spike gets shipped as a win.

The math is straightforward and worth internalizing: required sample size scales with the inverse square of the MDE. Halving the effect you want to detect quadruples the traffic you need. That's why chasing small lifts on low-traffic surfaces is usually a bad trade — do the power calculation first, and if the test needs eight weeks to detect your MDE, either pick a bigger swing or a higher-traffic surface.

One nuance for ratio metrics (conversion rates, click-through rates): variance depends on the baseline rate, so low-baseline metrics need dramatically more traffic. A checkout flow converting at 2% needs roughly 25x the sample of a page converting at 50% to detect the same relative lift. This surprises people every time.

Guardrail metrics: what you watch while the test runs

Your primary metric tells you if the change worked. Guardrail metrics tell you if it broke something else. Every test should declare, up front, a small set of guardrails: overall conversion or revenue, latency/page-load, error rates, and support-contact rate at minimum. The set depends on the surface, but the principle doesn't: no primary-metric win ships if a guardrail regresses.

The classic failure: a variant lifts click-through by making the button bigger and more aggressive, while quietly increasing accidental clicks, downstream drop-off, and support tickets. The primary metric says ship; the guardrails say the "win" is user frustration. I've seen this exact pattern more than once, and the guardrails are the only thing that caught it.

Treat guardrail breaches like SRM: investigate, don't rationalize. "It's probably noise" is the sentence that precedes every postmortem.

When to stop a test

Stop when one of these is true: you hit the pre-committed sample size (then decide by the rules you set before launch); a guardrail breaches badly enough to risk user harm (stop immediately, user safety overrides statistical purity); or SRM is detected (stop, fix, restart).

Don't stop because the result "looks stable" — stability at an underpowered sample is an illusion. Don't extend a test because you didn't like the result — extending until significance appears is peeking with extra steps, and it inflates false positives exactly the same way. And don't let tests run indefinitely "to gather more data": indefinite tests accumulate novelty decay, seasonality shifts, and overlapping experiments that muddy the read. A test should have a planned end date, and someone should own enforcing it.

The decision rule itself should be pre-registered: what p-value, what minimum effect size, and what guardrail conditions constitute a ship. Write it down before launch. The point isn't bureaucracy — it's that your future self, staring at a p=0.06 on a feature you personally built, cannot be trusted to set the bar objectively.

The bottom line

Good experimentation is mostly discipline, not math: pre-commit to sample sizes and decision rules, never peek at the primary metric mid-flight, monitor SRM and guardrails continuously, run long enough to outlast novelty, and power the test for an effect size you actually care about before it starts. The statistics are the easy part — libraries handle them. The hard part is building a team culture where "the dashboard looks good, let's ship it early" is recognized as what it is: the most expensive sentence in product development.

05 Beyond A/B: Multi-Armed Bandits for Continuous Optimization Thompson sampling mechanics, the regret-vs-statistical-power trade-off, and the failure modes that kill bandits in production. Read articleClose article

A fixed-horizon A/B test tells you which arm was better last month. A bandit tells you which arm is better right now — and routes traffic there while it's still finding out. Here are the machinery, trade-offs, and failure modes that matter.

Why A/B tests are an expensive way to learn

The standard A/B test has a hidden cost: while it runs, half your traffic goes to the thing you increasingly suspect is worse. On a high-traffic surface, that "learning tax" can dwarf the eventual lift. A/B is also fixed-horizon — ship the winner, and your knowledge starts decaying the day you ship.

Multi-armed bandits (MAB) exist for exactly this discomfort. Instead of learn, then earn, they earn while learning: every impression both exploits what you know and explores what you don't. The price is rigor — you trade statistical power for regret minimization, i.e., fewer wasted impressions along the way.

The two workhorses: epsilon-greedy and Thompson sampling

Epsilon-greedy is the algorithm you can implement in twenty minutes and explain in two. With probability ε you explore (pick a random arm), with probability 1−ε you exploit (pick the best observed mean). Teams typically start with ε around 0.1–0.2 and decay it over time.

It's honest and easy to reason about, but it explores uniformly: an arm that's been clearly terrible for a month still gets its ε/n share forever. With many arms — headlines, subject lines, notification copy — that uniform tax adds up.

Thompson sampling is smarter about where it explores. For each arm, you keep a posterior distribution over its true conversion rate (typically Beta(α, β) for binary outcomes, seeded with something like α=1, β=1). Each round, you sample one draw from each arm's posterior and play the arm with the highest draw:

for each arm a:
    sample θ_a ~ Beta(α_a, β_a)
play arm argmax(θ_a)
on reward r ∈ {0,1}: α_a += r;  β_a += (1 − r)

The elegance: an arm with a wide, uncertain posterior will occasionally produce a high draw and get explored; an arm that's clearly bad has a tight low posterior and almost never wins the draw. Exploration becomes proportional to uncertainty, which is exactly where it should go. Those Beta posteriors are also a dashboard: at any moment you can compute P(arm B beats arm A) by sampling from them — management gets a readable "we're 92% confident in the leader" without a statistics refresher.

Thompson sampling outperforms epsilon-greedy on regret in almost every realistic simulator, and it's barely more code. One caveat: Thompson sampling is not great at inference. It converges traffic to the winner so aggressively that losing arms end up with tiny samples, and your estimate of "how much worse was arm C?" is garbage. If you need that counterfactual, deliberately force exploration or run a proper randomized test.

When bandits beat fixed-horizon tests

Bandits win when three conditions hold simultaneously:

  1. High traffic relative to arm count. Bandits need feedback to adapt. If you get 200 conversions a day and you're testing 40 variants, you're not bandit-ing — you're gambling with extra steps. Rule of thumb: each arm needs meaningful feedback within days, not months.
  2. The cost of being wrong is continuous, not one-shot. Homepage hero copy, notification send-time, search ranking blends — surfaces where a suboptimal arm burns money every hour it runs.
  3. The environment is plausibly non-stationary. This quarter's winning headline may not win next quarter. A/B tests answer "which was better during the test window"; bandits answer "which is better now, given everything seen so far" — as long as you handle recency correctly (more on that below).

Where bandits don't beat A/B: big-bet product changes (you can't "bandit" a pricing page redesign — the cost of flipping users back and forth is absurd), anything needing causal inference for compliance, and low-traffic experiments where significance math was already doing all the work.

The regret vs. statistical power trade-off

Cumulative regret — the total reward lost to not always playing the best arm — is the metric bandits optimize. Statistical power is what A/B tests optimize: the ability to detect a true difference of size δ with probability 1−β.

These are in tension by construction. A bandit that converges fast starves the losers of samples, so you lose the ability to say anything rigorous about them. Worse, the sample you do collect is adaptively biased — allocation depended on past outcomes — so naive confidence intervals on bandit data are invalid. Corrections exist (inverse propensity weighting, always-valid confidence sequences), but they're fiddly in a streaming pipeline and stakeholders won't intuitively trust them.

My practical stance: decide up front which currency you're being paid in. If the org needs a defensible "we shipped X because it lifted conversion by 3.2% ± 0.8%" — run a randomized test. If the org needs "we stopped bleeding revenue on this surface" — run the bandit, report posterior probabilities, and don't pretend you have confidence intervals you don't. The teams that get burned run a bandit and then write an A/B-style readout from its data.

Cold start and non-stationarity

Every bandit begins in the worst possible state: zero information, maximum regret. Two things help. The Beta(1,1) uniform prior is a fine default, but if you have historical data for the same surface — last month's conversion rates — seed the posteriors with it. An arm that historically converts at 5% should start near Beta(50, 950), not Beta(1,1). This is the highest-leverage knob on early regret — but downweight history with a decay, or stale priors in a non-stationary world become actively harmful.

Non-stationarity needs deliberate forgetting. A plain Beta-posterior bandit has perfect memory: an arm that won in January carries that evidence into June. The standard fixes are sliding windows (keep only the last N observations), exponential discounting (weight past observations by γᵏ, γ ≈ 0.99–0.999 per update), or a drift detector that resets posteriors on regime change. I reach for exponential discounting first: one parameter, no sharp edges, degrades gracefully.

Where bandits fail

I've watched bandits fail in four characteristic ways:

1. Delayed rewards. Bandits assume the reward for an action arrives quickly. In many real systems it doesn't: a notification's effect on 7-day retention takes 7 days; a pricing test's effect on churn takes months. By the time the feedback arrives, the bandit has already reallocated. Fixes exist — surrogate short-term rewards, reward imputation, batched update schedules — but each injects assumptions. If the delay exceeds your adaptation timescale, a fixed-horizon test whose analysis window you can reason about is probably the better tool.

2. Interference and network effects. Bandits assume each play is independent. In marketplaces, social products, and shared inventory systems, giving arm A to one user changes the experience for another. The arm-level reward you observe is contaminated. This is the stable unit treatment value assumption breaking, and no bandit variant fixes it — you need cluster randomization or marketplace-specific designs.

3. Reward hacking through the metric. The bandit optimizes whatever reward function you give it, ruthlessly. Feed it click-through and it will learn to serve rage-bait. I've seen a notification bandit converge on the most annoying-but-openable copy within days. The lesson bears repeating in bandit form: the regret you're minimizing is defined by your reward signal, and the bandit will find every degenerate corner of it faster than a human team would.

4. Operational fragility. A bandit is a live control loop on a revenue surface — monitor it like one: allocation shares per arm, posterior entropy, reward ingestion rate. The scariest failure I've seen was a logging bug that silently dropped negative rewards, so the bandit converges to the worst arm because it never saw its failures. Put the bandit's inputs on the same alerting footing as the surface it optimizes.

The bottom line

A/B tests are measurement instruments; bandits are control systems. Choose measurement when you need to know, control when you need to earn. Thompson sampling with sensible priors and discounted history is a fine default for the control case — but it won't give you the statistical story an A/B test would, it won't survive delayed rewards without help, and it will optimize exactly the metric you hand it, pathologies included.

06 The Growth Engineering Stack: Instrument First, Optimize Second Event taxonomy as schema contract, deterministic-first identity resolution, the warehouse as source of truth, and the failure modes of bad instrumentation. Read articleClose article

Every growth team eventually learns the same lesson the painful way: the funnel analysis was wrong because the events were wrong. Here's how to build an instrumentation layer that doesn't move under you.

The sequencing argument

The conventional growth playbook goes: form hypothesis, run experiment, measure, iterate. It quietly assumes the "measure" step is solved. In my experience it's the step that's broken most often and fixed last, which is exactly backwards.

Consider what happens when instrumentation is an afterthought. The signup funnel shows a 40% drop between step 2 and step 3. The team spends three weeks testing button copy before someone notices the step-3 event fires on page load rather than on successful completion — it counts everyone who bounced mid-render. The "40% drop" was never real. Multiply this across every team and you get the ambient condition of most growth orgs: decisions made at full speed on data nobody trusts — so the data is ignored and decisions get made on instinct anyway.

Instrumentation-first flips the sequence: before you optimize anything, you make the measurement layer boring. Events are defined, documented, reviewed, and tested like API contracts. Only then does experimentation earn its keep. Slower for the first quarter, dramatically faster forever after — you stop paying the tax of re-litigating every number.

Event taxonomy: the boring design that matters most

A growth stack lives or dies on its event taxonomy. Getting it right is unglamorous; getting it wrong is a slow-motion disaster. The principles that survive contact with real teams:

Name events as past-tense, object-verb pairs. Checkout_Completed, not checkout or PurchaseButtonClicked. The event should describe the thing that happened in the world, not the UI interaction that may or may not have caused it. This matters because UIs get redesigned constantly; the underlying business moment — the user paid — is stable. When the name encodes the UI, every redesign orphans your history.

Properties are the schema; treat them as such. Every event carries a fixed set of typed properties with documented semantics: amount is in cents, always; currency is ISO 4217; source is an enum, not free text. The failure mode here is the "properties junk drawer" — engineers stuffing misc_info strings and ad-hoc fields into events because adding a proper property required a schema review. Six months later, analysts are parsing misc_info with regexes and the warehouse is a swamp. The fix is boring process: a schema registry enforced by a CI check on the tracking plan, versioned changes, and a rule that no event ships without its property contract reviewed.

Version deliberately, not accidentally. There are two honest approaches: version the event (Checkout_Completed_v2) or keep the name and evolve properties with additive-only changes. Additive-only is cleaner to query but requires discipline about never renaming or retyping a property. Versioned events are messier downstream (every query needs WHERE event IN (...)) but honest about breaks. What kills you is the third approach — silently changing semantics while keeping the name — which is what happens without a process. Pick one of the first two and enforce it in code review.

Identity resolution: the hardest problem you'll pretend is solved

Every interesting growth question is a cross-session, cross-device question: did the person who saw the email on their phone complete signup on their laptop? Answering it requires stitching anonymous and identified identities into a single user graph.

The machinery has two halves. Deterministic stitching — joining on a stable identifier like a logged-in user ID or an email hash — is reliable and should be the backbone. Probabilistic stitching — fingerprinting, device graphs, heuristic session merging — fills the gaps but introduces false merges, which are poison: two users merged into one corrupts every downstream metric in ways nearly impossible to detect from the metrics alone.

My strong opinion: build the deterministic path first and make it excellent — consistent ID assignment at signup, login, and every auth-adjacent moment; a single user_id in every event; session IDs that survive app restarts. Then measure your identity coverage — the fraction of events carrying a resolved user_id — as a first-class data quality metric, and only reach for probabilistic methods when deterministic coverage is genuinely insufficient. Most teams I see do this in reverse: they buy a fancy identity product while their own login flow emits three different IDs for the same user.

The warehouse as source of truth

Product analytics tools are wonderful for exploration and terrible as systems of record. Their definitions drift (one PM's "active user" is not another's), their data retention is finite, and their export formats change. The warehouse — your own copy of the raw event stream, in your own storage, under your own schema — is the only thing you fully control.

The practical shape: clients and servers emit events to a collector, which lands raw payloads in immutable storage (object storage, append-only), and a pipeline normalizes them into the warehouse with deduplication, schema validation, and late-arriving-data handling. Derived marts (funnel tables, user-level aggregates, experiment assignment tables) are built on top of the raw layer, never instead of it. When — not if — you discover a definition was wrong, you rebuild from raw. Teams that skip the raw layer and go straight to aggregated tables discover, usually during a board-meeting argument, that they can't answer "what did we actually measure?" for anything older than the current definition.

There's also a less obvious reason: the warehouse is where growth engineering meets the rest of the company. Finance reconciles revenue against it. Data science trains models on it. Marketing attribution reads from it. If the warehouse is an afterthought, every consumer builds their own shadow pipeline, and you end up with five definitions of revenue and a quarterly reconciliation ritual. Instrument once, into the warehouse, and let everyone read from the same raw material.

Failure modes I've seen (so you don't have to)

Event drift. The taxonomy is documented, and then the product ships anyway. New events appear with ad-hoc names; old events silently change meaning after a refactor. The defense is automated, not cultural: a tracking plan as code, CI checks that fail the build when an event lacks a schema entry, and a weekly automated diff of "events seen in production vs. events in the tracking plan." Drift you can see is drift you can fix.

PII leakage into the event stream. Somewhere in every org, a well-meaning engineer has put an email address, a full name, or a precise location into an event property "for debugging." Once it's in the raw stream, it's in every downstream copy, every derived table, every export — and now your retention and deletion story is a lie. The fix is a denylist enforced at the collector (drop or hash known PII patterns before they land) plus making the safe path the easy path: provide a setUserProperties API with explicit PII handling so engineers never need the event stream for it. Audit quarterly; you will find something.

Duplicate events. Double-firing from retry logic, events sent from both client and server "for redundancy," page-reload replays — duplicates are the most common silent corruptor of funnel metrics I've encountered. The defense is idempotency: every event carries a client-generated UUID, and the pipeline dedupes on it. A useful diagnostic: the ratio of distinct event UUIDs to total rows should be ~1.0; alert when it isn't.

Sampling without telling anyone. At scale, someone will suggest sampling the event stream to save cost — a reasonable idea, except when applied silently to the tables the growth team queries. A 10% sample is fine for trends and fatal for rare-event funnels. If you sample, sample deterministically (hash-based on user_id so the sample is stable and joinable), document the rate next to every derived table, and never sample the raw layer.

The payoff

Instrumentation-first is a hard sell because its output is invisible: it's the absence of arguments about numbers, the dashboard everyone quietly trusts, the experiment readout nobody re-litigates. But I've watched teams go from "we can't trust the funnel" to shipping two experiments a week on the same headcount, and the entire difference was six weeks of taxonomy, identity, and pipeline work that nobody demoed.

Optimize second. Instrument first. The compounding starts immediately.

07 The Modern Marketing Data Stack, Explained for Engineers What CDPs actually do, reverse ETL sync semantics, consent as an engineering problem, and where the stack breaks. Read articleClose article

CDPs, reverse ETL, consent platforms, identity graphs — the marketing data stack is a distributed systems problem wearing a costume. Here's what each piece actually does, where it's genuinely hard, and where it breaks.

Why engineers should care

If your company has a marketing team, you're already operating this stack. The signup events your app emits feed the CDP; the warehouse tables you maintain power the ad audiences; the consent banner your frontend renders gates all of it. When the stack breaks, the symptoms land on engineering — "why did 40,000 users get the win-back email after they already converted?" is a data pipeline bug with a marketing face.

The CDP: what it actually does vs. the marketing

Strip away the vendor decks and a Customer Data Platform does four concrete things:

  1. Ingest event and profile data from your app, website, and backend.
  2. Resolve identity — stitch the anonymous cookie, device ID, and logged-in user into one profile.
  3. Compute traits on profiles: "days since last purchase," churn scores, segment membership — a streaming feature store scoped to marketing.
  4. Activate — push profiles, traits, and segments out to email, push, ads, and personalization tools.

That's it. The "single view of the customer" is true the way a data lake is "a single view of your data" — its quality is entirely a function of what you put in and how you govern it. An engineer evaluating a CDP should ask three questions: how does identity resolution work, what's the latency from event-in to segment-out, and can I get my data back out? A CDP you can't exit is a roach motel for your customer data.

Reverse ETL: the warehouse strikes back

Reverse ETL inverts the usual flow: instead of data leaving your warehouse into marketing tools, the warehouse becomes the source of truth and syncs audiences, traits, and predictions outward to ad platforms, email tools, and sales systems.

Why this matters: the warehouse already holds the cleanest customer data — finance reconciles against it, models train on it, definitions are versioned there. Reverse ETL lets marketing act on it without building a parallel, drifting copy in every SaaS tool. A segment like "high-value users at risk of churning" can be defined once, in SQL, tested in CI, and synced to five destinations.

The engineering is in the sync semantics: full refresh or diff? What happens when the destination API rate-limits you mid-sync — retry, skip, or partial-write? I've seen a full-refresh sync briefly empty an ad-platform audience because the tool interpreted "replace audience" literally. Same discipline as any distributed sync: idempotent writes, diff-based updates where supported, and a dry-run mode.

Consent management and its engineering implications

Consent — GDPR, CCPA/CPRA, and the expanding patchwork — is usually filed under "legal," but it's an engineering problem: every downstream use of personal data must be gated on current consent state, and consent changes must propagate everywhere, quickly.

This creates three engineering obligations most teams underestimate:

  • Consent as a first-class event. Consent changes should flow through the same pipeline as everything else — timestamped, immutable, replayable. If consent state lives only in the CMP's dashboard, you can't reconstruct "what were we allowed to do with this user on March 14th?" — exactly the question a regulator asks.
  • Propagation latency budgets. Measure end to end how long an opt-out takes to reach every destination. I've seen stacks where the banner updated instantly and the ad-platform sync ran on a 24-hour batch — a full day of unlawful processing, invisible except in the logs. Set an SLO on consent propagation like any critical path.
  • The deletion cascade. A deletion request has to fan out to the warehouse, the CDP, every activated destination, backups, and derived tables. If your raw event layer is append-only, you need a designed mechanism — tombstoning, crypto-shredding, or partitioned deletion — not a manual ticket queue.

The non-obvious insight: consent infrastructure is identity infrastructure. You can't gate "this user's data" on consent unless you know which data is "this user's." Teams that treat consent as a banner plugin discover during an audit that it's the same problem as identity resolution.

Real-time activation vs. batch: pick your latency honestly

The stack splits into two latency regimes, and conflating them causes most architectural arguments:

Batch (hours to a day): warehouse-driven segments, overnight model scores, scheduled campaigns. Cheap, reliable, debuggable — and completely adequate for most marketing use cases.

Real-time (seconds to minutes): event-triggered messages, on-site personalization, fraud-adjacent gating. Requires streaming infrastructure, exactly-once-ish semantics, and on-call burden to match.

The failure mode is reaching for real-time because it sounds better. I've watched teams build a streaming pipeline for a "real-time" win-back email the business then scheduled to send once daily anyway. Real-time has a genuine cost curve — stream processing, state stores, late-data handling, 3 AM pages — and should be spent where latency changes outcomes: cart abandonment, transactional triggers, in-session personalization. For everything else, batch is not a compromise; it's the correct choice. Make the latency requirement explicit per use case and let the cheap regime win by default.

Identity graphs: the load-bearing assumption

Under the CDP sits the identity graph — the mapping between cookies, device IDs, emails, and user IDs that decides "these events belong to the same person." Everything downstream inherits its errors.

First, the graph is probabilistic at the edges: deterministic matches are trustworthy; probabilistic matches are guesses with confidence scores, and vendors set the threshold differently. A vendor boasting "98% identity resolution coverage" may be buying it with aggressive merging — meaning your "personalized" campaigns sometimes personalize for the wrong person. Ask for the precision/recall trade-off, not the coverage number.

Second, the graph degrades structurally over time. Third-party cookie deprecation, Apple's ATT opt-in rates, users with three devices and two emails — the raw material for identity resolution is getting thinner. The durable response is first-party identity: logged-in states, your own identifiers, value exchange that earns the login. The graph is load-bearing, and the load is getting heavier while the foundation gets lighter.

Where the stack breaks

Schema drift across the seams. The app team renames an event property; the CDP mapping still references the old name; the segment "users who did X" silently goes to zero; the campaign sends to nobody for a week before anyone notices. Every seam — app to CDP, CDP to warehouse, warehouse to reverse ETL, reverse ETL to destinations — is a schema contract that can drift. Contract testing across seams (a nightly job asserting the critical fields are non-null in the last 24h of synced data) catches most of it.

Consent propagation latency, covered above — the gap between the banner and the last destination.

Segment staleness. A segment computed from yesterday's warehouse sync is used to suppress today's "come back!" push — and the user who converted this morning gets the win-back message anyway. Every segment needs a documented freshness SLA.

Destination API fragility. Ad platform APIs change, rate-limit, and occasionally just break. Treat destinations as unreliable third parties — retries, dead-letter queues, per-destination sync health alerting. A silently failing sync is a budget being spent on a stale audience.

Attribution theater. Last-click, multi-touch, modeled, incrementality-tested — the stack will happily produce any attribution number you ask for, and different tools disagree on the same data. My stance: the only attribution number that means anything comes from a properly designed incrementality test; everything else is an allocation heuristic wearing a lab coat. Build the stack so holdout-based incrementality tests are easy, and label the heuristics as heuristics.

The engineer's takeaway

The marketing data stack is a distributed system with marketing-shaped requirements: identity resolution is entity resolution, consent propagation is a consistency problem, reverse ETL is change data capture, real-time activation is stream processing. Bring your systems instincts — contracts at seams, idempotency, freshness SLAs, graceful degradation — and the stack is less mysterious than its vendors want it to be.

08 Exactly-once is a lie (mostly): delivery semantics in distributed systems, in practice Why exactly-once is unachievable in the general case, what systems actually guarantee, and at-least-once + idempotency as the real workhorse. Read articleClose article

Every messaging system eventually advertises an "exactly-once" mode. I've enabled those flags, and I've cleaned up the duplicates they were supposed to prevent. This is the gap between what "exactly-once" means in a vendor whitepaper and what it costs to actually get it — and the pattern that carries almost all production traffic anyway.

Why exactly-once is impossible in the general case

The argument is older than most of the systems we run. Two generals need to agree on a plan but can only communicate through messengers who might be captured. Any protocol that ends with "I received your confirmation" leaves one side uncertain: did the last message arrive? No finite sequence of acknowledgements resolves this, because the final message in any protocol is always unacknowledged.

Delivery is the same problem in a different uniform. A producer sends a message and needs to know it was processed, not just received. The acknowledgement can be lost, the network can partition between "processed" and "ack sent", and the consumer can crash after applying the effect but before checkpointing. In that window the producer cannot distinguish "never processed" from "processed, ack lost" — so it retries, and the consumer sees the message twice.

This isn't a limitation of engineering effort. Over an unreliable network you can choose at-least-once (retry until acknowledged, accept duplicates) or at-most-once (never retry, accept loss). "Exactly-once" — deliver once and never lose — is unachievable as a network property. Everything marketed under that name is a narrower claim in a general-sounding label. The interesting question is always: narrower how?

What systems actually mean by "exactly-once"

Kafka transactions. The Kafka story is the idempotent producer (per-partition sequence numbers; the broker rejects duplicates) plus the transactional API: consume from topic A, process, produce to topic B, and commit the consumed offsets as one atomic unit. A consumer on isolation.level=read_committed never sees aborted writes.

The boundaries are where production bites. First, the guarantee covers the log, not your side effects. If processing calls an external API — charges a card, sends an email — and the transaction then aborts, the card was still charged. Kafka cannot roll back the world. Second, zombie fencing (transactional.id + epoch) stops a stale producer from writing after failover, but it doesn't help when your consumer crashes after processing and before committing offsets: the next poll redelivers, and the processing ran twice. Third, the guarantee holds only if every consumer in the chain cooperates — one read_uncommitted consumer downstream and the semantics silently degrade. I've seen an exactly-once pipeline where one legacy consumer nobody owned was reading uncommitted data. The flag was on; the guarantee was off.

Flink checkpoints. Flink snapshots operator state to durable storage and, on recovery, rewinds sources to the last checkpoint. Internal state is exactly-once: your running aggregations come back correct. End-to-end exactly-once additionally requires two-phase-commit sinks — the sink joins the checkpoint protocol and commits only when the checkpoint completes. For Kafka-to-Kafka this works. For sinks without 2PC support (most of them: a REST endpoint, a database without XA, an email service), Flink quietly gives you at-least-once at the edges with exactly-once state in the middle. The docs are honest about this; pipeline diagrams rarely are.

The pattern: every honest exactly-once claim is scoped to a transactional boundary — the set of systems participating in the same commit protocol. Inside the boundary: exactly-once. Outside it: hope. And the boundary almost never includes the thing you actually care about, which is the business side effect.

The workhorse: at-least-once plus idempotency

Since the network won't give you exactly-once, production systems take the guarantee the network can give and make duplicates harmless: the effect is applied once even if the message arrives three times. This is the pattern behind essentially all payment processing, webhook handling, and queue consumers at scale — and it's more robust than any broker flag because it survives broker changes, consumer rewrites, and the 2 AM failover nobody predicted.

The mechanism is a deduplication record keyed by an idempotency key, stored transactionally with the effect:

def handle_charge(msg):
    key = msg.idempotency_key   # generated once by the sender, reused across retries
    with db.transaction():
        if db.exists("processed_keys", key):
            return db.get_result(key)  # duplicate: return the original outcome
        result = charge_card(msg.amount, msg.card_token)
        db.insert("processed_keys", key, result)
        return result

Two details make or break this. The dedup check and the effect must commit atomically — the check in the same transaction as the write, or you re-open the crash window. And on a duplicate you must return the original result, not just skip: the first attempt's response may have been lost, and the retrying caller needs the answer, not silence.

Idempotency key design is where I've seen the most quiet failures:

  • Generate once, at the edge, reuse across retries. The key is created by the originator before the first attempt and attached to every retry. If each retry mints a new key, your dedup table is decorative.
  • Prefer natural keys over UUIDs. If the operation has a business-unique identifier — order ID, payment reference, (account, statement_date) — use it. Natural keys deduplicate across different code paths converging on the same operation, not just retries of one call.
  • Key the dedup record at the right granularity. Dedup on (operation, key), not just key — the same key hitting two different endpoints must not collide. And distinguish "retry of charge $50" from "a new $50 charge for a different order."
  • Make keys deterministic when the sender can't store them. A producer that can't keep per-request state can derive the key from stable inputs: HMAC(event_type || source_id || source_timestamp). Deterministic derivation survives sender crashes and replays from cold logs.

Dedup windows are the part nobody budgets for. A dedup table grows without bound, so in practice you keep it for a TTL — 24 hours to 30 days, depending on your maximum retry horizon. The window must exceed the longest time a duplicate can plausibly arrive: your retry backoff ceiling, your queue's max retention, your consumer's worst-case lag during an incident. I've seen duplicates arrive days after the original during a major backlog drain, sailing past a 24-hour dedup window. Size the window for the incident, not the sunny day — that sets your storage bill.

Where duplicates still leak through

Even well-built idempotent consumers have seams. The ones I've actually been paged for:

  1. Retries racing in-flight requests. The client times out at 5s and retries while the first request is still executing; both pass the dedup check before either inserts. Fix it with a unique constraint that makes the second insert fail atomically, or insert the key in a "processing" state before executing and have the racing request wait or return 409.
  2. Non-idempotent steps inside the handler. The dedup record covers your handler, not the five downstream calls it makes. If your handler charges the card (keyed) and then sends a confirmation email (not keyed), a crash between them double-sends the email. Key every side effect, or restructure so each effect is its own idempotent handler.
  3. Dedup store failure modes. The dedup table is now on your critical path — its outage either blocks all processing or gets bypassed "temporarily," and every retry becomes a duplicate. Treat it with the same durability as the data it protects: replicated, backed up, never skipped on the happy path.
  4. Clock and ordering assumptions. Deduplicating "the latest" of several versions assumes you can order versions correctly. Across partitions and retries, arrival order is not causal order; source-assigned sequence numbers beat wall-clock timestamps.
  5. Multi-region and cross-system replays. Backfills, region failovers, and "replay the last week from the archive" generate duplicates outside your retry window and sometimes outside your key scheme (replayed events may carry new keys). Decide up front whether replays must be idempotent; if yes, the key must be derivable from the event itself, not assigned at ingestion.

None of this is an argument against Kafka transactions or Flink checkpoints — use them; they genuinely shrink the surface where duplicates form. It's an argument about where the guarantee lives. The broker flag narrows the window; idempotent receivers close it. Design for at-least-once, make every effect idempotent, size your dedup windows for your worst incident, and audit every side effect for its own key. That is what "exactly-once in practice" has ever meant.

09 Temporal vs Airflow vs Conductor: what benchmarking orchestration engines taught me Architectural trade-offs in durability models, developer experience, and operational cost, drawn from building a reproducible benchmark harness (SoftwareX, DOI 10.1016/j.softx.2026.102998). Read articleClose article

I spent the better part of a research project building a reproducible benchmark harness that runs the same workflow shapes against Temporal, Netflix Conductor, and Apache Airflow, and published the work as sole author in SoftwareX (Elsevier; DOI 10.1016/j.softx.2026.102998). The point of the paper was a fair, repeatable comparison harness. But the durable lessons weren't the measurements — they were architectural. Building the harness forced me to understand each engine's durability model, failure semantics, and operational shape deeply enough to make the comparison honest, and that exercise taught me more about choosing orchestration than any feature matrix.

One note on scope: I'm deliberately not reciting benchmark figures — they're meaningful only with the harness config, workload shape, and tuning attached. What follows is qualitative: the trade-offs that decide which engine fits which problem.

The durability models are the whole story

Everything else about these engines is downstream of one decision: where does workflow state live, and who advances it?

Temporal is event-sourced. A workflow execution is a history of events persisted by the Temporal server; workers run workflow code, and on recovery the code replays from history to rebuild state before continuing. Because state is a replayable log rather than a mutable row, a workflow can sleep for a year, survive every worker dying, and resume exactly where it left off. The cost is the replay constraint: workflow code must be deterministic, since it will be re-executed. Random numbers, wall-clock reads, unordered-map iteration — all must go through Temporal's APIs, which record the values in history. This is a genuine tax, and it's the first thing that breaks when someone ports "normal" code into a workflow.

Airflow polls a database. The scheduler loops over the metadata DB, DAG definitions produce task instances with states, and workers claim and run them. Durability is row state in Postgres/MySQL: queued, running, success, failed. Simple and legible — you can read the DB and know exactly what's happening — but the workflow itself has no durable memory beyond task states and XComs. There is no "resume from line 47." Recovery is per-task: retry the task, from the start. For ETL-shaped work this is exactly right; for a long-running business process with branching logic and timers, it's a poor fit, and no operator cleverness changes the model.

Conductor decouples orchestration from execution. The Conductor server holds workflow definitions (JSON) and execution state in its datastore, but tasks run on external workers that poll task queues. The engine decides what's ready; your services decide how to run it. The orchestrator never runs your code, which makes it polyglot-friendly and keeps the engine horizontally scalable. The trade-off: everything the engine knows about a task is what the worker reports back. Rich in-workflow logic — data-dependent branching, long sleeps, complex compensation — lives awkwardly in JSON definitions, so real logic migrates into the workers and the "workflow" becomes a coordination skeleton.

Building the harness made this concrete: the same logical workload had to be expressed three different ways, and the translation was the benchmark.

Developer experience: three different contracts

Temporal: workflows as code, with rules. You write what looks like ordinary code — function calls, loops, try/catch — and the engine makes it durable. The DX win for complex logic is enormous: a workflow reads like the business process it implements. The contract: determinism constraints, the replay model, and versioning discipline when workflow code changes mid-flight. Teams that treat workflows like regular code get burned; teams that treat them like database migrations do fine.

Airflow: DAGs as Python, schedule-first. The mental model is a batch pipeline: define tasks, wire dependencies, schedule the DAG. Every data engineer already thinks this way, which is why Airflow won its category. But the Python is configuration that looks like code — top-level DAG-file code runs on every scheduler heartbeat, which surprises newcomers (no heavy imports, no I/O at module scope), and dynamic branching fights the static-DAG grain. The contract: your logic must fit "tasks with dependencies on a schedule."

Conductor: JSON definitions, code in workers. Definitions are JSON (or builders that produce JSON); all real logic lives in worker services you operate. Liberating for polyglot shops — Go, Python, and Java workers can serve one workflow — and the right shape when tasks are really "call this microservice and wait." The contract: you're operating a distributed system of your own workers, with their own deploys, scaling, and failure stories, and the definition language will never be as expressive as code.

What you run yourself

The part of the comparison that gets hand-waved in feature matrices and dominates real adoption decisions.

Temporal: the server cluster (frontend/history/matching services), a visibility store (Elasticsearch), and primary persistence (Cassandra/MySQL/Postgres). A genuinely distributed system to operate — fair, since it provides genuinely strong guarantees. The managed offering exists because many teams don't want this on their plate.

Airflow: scheduler, webserver, metadata DB, and an executor backend (Celery workers plus a broker, or KubernetesExecutor…). The component count is high and failure modes spread across them: scheduler stalls, worker OOMs, DB connection exhaustion, zombie tasks. Airflow's operational reputation is earned.

Conductor: the server (API + orchestration), its datastores, and — critically — your worker fleet. Conductor-the-engine is relatively light; Conductor-the-system includes everything your workers need. Teams sometimes adopt it expecting "lightweight orchestration" and discover they've signed up to operate a fleet of services anyway. If those services already exist as your microservices, that's free. If not, it's the real cost.

The lesson from running all three under a harness: measure operational cost in components you must understand when paged, not container count. All three page you; they just page you about different things.

Failure recovery semantics

Temporal recovers workflows, not tasks. Kill every worker mid-execution; new workers replay histories and continue. The unit of recovery is the workflow execution, and completed activities are never redone. The strongest story of the three for long-running processes — and why Temporal fits sagas and human-in-the-loop flows: the engine remembers everything.

Airflow recovers tasks. A failed task retries per its retry policy; downstream tasks wait. If the scheduler dies, the DB still holds task states and a new scheduler picks them up. But partial progress within a task is lost unless the task checkpoints externally — Airflow has no replay. Idempotent tasks aren't a best practice here; they're load-bearing.

Conductor recovers by re-queueing. A worker that dies mid-task leaves the task IN_PROGRESS until a timeout fires, then it's re-scheduled to another worker. Simple and robust — but the timeout is the detection mechanism. Too long and failures linger; too short and slow tasks get double-executed.

Temporal remembers the most and asks the most; Airflow and Conductor remember task states and ask you to make tasks atomic and retryable.

When I'd pick each

After building the harness and living with all three, my selection logic is about the shape of the work:

  • Temporal, when the workflow is a long-running process with real logic — branching, timers, signals, human steps, compensation across services. Order fulfillment sagas spanning payment, inventory, and shipping; onboarding flows that run for days; anything where "resume exactly where we were" is the requirement. The determinism tax pays for itself the first time a year-long workflow survives an incident without manual repair.
  • Airflow, when the work is scheduled data pipelines. Batch ETL/ELT, periodic reporting, ML training on a cadence. The schedule-first model, the operator ecosystem, and the data-engineering community are compounding advantages no general-purpose orchestrator matches here. Don't fight it into being a microservice orchestrator.
  • Conductor, when you have a fleet of services to coordinate. Polyglot environments, existing microservices that should stay dumb and independently deployable, event-driven flows that need a conductor (pun intended). The thinnest layer that still gives you visibility and control over cross-service flows.

And the meta-lesson from the benchmarking project itself: reproducibility is the feature. The reason the paper exists as a harness rather than a table of numbers is that orchestration performance without the workload definition, tuning, and failure injection described is marketing. Any engine comparison worth trusting ships the harness, pins the versions, and describes the workload shapes — because the shapes decide the winner before the first measurement runs. If you're evaluating orchestrators, build the smallest version of your workflow in each candidate and fail it on purpose. How it recovers tells you more than any benchmark I could publish.

10 Serving LLMs at scale: batching, caching, and cost control Continuous batching, the KV-cache concurrency ceiling, prefix vs. semantic caching, small/large model routing, and honest cost-per-token accounting. Read articleClose article

Running an LLM in a demo and running one for millions of requests are barely the same discipline. Inference is where your product's token economics get decided — and most of the cost isn't the model, it's how you serve it. Here's what actually moves the needle.

Continuous batching beats static batching

The first LLM serving systems used static batching: collect N requests, run prefill for all of them, decode until every request finishes, then start over. The problem is the straggler effect — if request A needs 10 output tokens and request B needs 500, A sits idle for B's entire decode while still occupying its slot and KV cache. GPU utilization collapses toward the longest request's tail latency.

Continuous (iteration-level) batching — the approach pioneered in Orca, now standard in vLLM, TGI, and TensorRT-LLM — schedules at the granularity of a single decode step. As soon as request A finishes, its slots are freed and a newly arrived request takes its place — mid-batch, mid-generation. The GPU stays saturated because the batch is constantly being repacked.

The non-obvious consequence: server behavior changes fundamentally under load. At low QPS, per-request latency is dominated by prefill. As QPS rises, effective batch size grows — latency degrades gracefully while throughput climbs — until decode iterations slow because attention over a very large batch saturates the GPU. That saturation point is where your SLO breaks, and it's the single most important number to find empirically for your workload. Think of capacity as "tokens per second at p99 latency X" — tokens are the currency the hardware understands.

The KV cache is your real concurrency ceiling

Autoregressive decoding needs the keys and values of every previously generated token, and for large models this cache dwarfs the model weights — roughly 1–2 MB per token for a 70B-class model. A request generating 4k tokens holds several GB of GPU memory for its lifetime.

This is the binding constraint on concurrency, not FLOPs. If your GPU has 80 GB and weights take 140 GB across a tensor-parallel setup, the leftover memory divided by per-request KV usage is your max concurrent requests, full stop. Techniques like PagedAttention (virtual memory for KV blocks) don't change the ceiling — they remove fragmentation, so you actually reach the ceiling instead of hitting it at 40% utilization.

Practical implications teams miss:

  • Max sequence length is a capacity decision, not a model capability. Setting max_tokens=8192 when your p99 is 400 tokens reserves memory you never use. Admission control reserves per the configured max, so a generous max quietly halves your concurrency. Set max tokens per route, per use case.
  • Quantization buys KV headroom too. Dropping weights to FP8/INT8 reduces the memory-bandwidth cost of every decode step, which is what decode is bottlenecked on. KV cache quantization can cut cache size in half with negligible quality impact on most tasks.
  • Disaggregated prefill and decode (different GPU pools) exists because their resource profiles differ: prefill is compute-bound and likes big batches; decode is memory-bandwidth-bound and likes steady streaming. When your prefill/decode ratio is extreme, mixing them on one pool means one phase always starves the other.

Caching: prefix caching is free money, semantic caching is a gamble

Two very different techniques share the word "cache." Conflating them causes bad designs.

Prefix caching stores the computed KV cache for a shared prompt prefix. If your system prompt plus few-shot examples are 2k tokens and every request starts with them, recomputing those 2k tokens on every request is pure waste — prefill is compute-bound, so skipping it directly raises throughput. With radix-tree-based cache management (vLLM's automatic prefix caching), any shared prefix across requests gets reused automatically. Hit rates in agentic workloads with fixed system prompts routinely exceed 50% — the closest thing to a free lunch in LLM serving. The catch: it only helps prefill compute, and only when requests actually share a prefix.

Semantic caching stores (embedding of prompt → response) pairs and serves near-duplicate queries without invoking the model at all. The hit-rate realities are harsher than the marketing suggests: on genuinely diverse user traffic, semantic hit rates are often 5–15%. The failure mode is quality, not performance — two prompts can be near-identical yet require different answers. Every semantic cache needs a similarity threshold tuned against your tolerance for wrong answers, plus an invalidation story for anything time-sensitive. I treat it as a cost optimization for high-volume, low-stakes queries, not a latency strategy — a cache miss followed by a full model call plus an embedding round trip is slower than no cache at all.

A rule of thumb that has held up: invest in prefix caching first (deterministic, no quality risk), then consider semantic caching only if you can measure duplicate traffic and can tolerate the occasional wrong-but-plausible answer.

Route between models, don't just scale one

The cheapest token is the one the small model generates. A well-designed serving tier routes queries across model sizes: a fast classifier or the small model itself handles the easy 70–80% (classification, extraction, short answers), and only the hard tail escalates to the large model. If an 8B model handles 75% of traffic at 1/10th the cost per token, blended cost drops by roughly 2/3.

The hard part is the router, and there are two honest approaches. A learned router (a small classifier trained on examples of "small model succeeds/fails") is cheap and works when your query distribution is stable. Self-verification (small model answers, then a verifier checks) is more robust to distribution shift but burns some of the savings. What doesn't work: keyword heuristics that nobody maintains. An unmaintained router silently degrades into "everything goes to the big model" — the most expensive configuration wearing the costume of an optimization.

The latency–throughput–cost triangle

You get to pick two, and the third is determined. The relationships are concrete, not philosophical:

  • Latency ↓, throughput fixed → cost ↑. Low TTFT means small batches and more GPUs per unit of traffic — interactive tiers cost 2–4x per token what batch tiers cost.
  • Throughput ↑, latency fixed → cost ↑. Scale out GPUs; past single-pool saturation you add load balancers, cross-region replication, and fleet operational tax.
  • Cost ↓ → latency or throughput suffers. Higher batching, smaller models, aggressive quantization — each trades quality or tail latency for dollars.

How to navigate it in practice: define the triangle per workload, not per company. Split traffic into tiers (interactive / nearline / batch), give each its own SLO and its own serving configuration, and let the batch tier run at maximum batch size with the cheapest acceptable model. The triangle stops being a dilemma once you stop forcing all traffic through one vertex.

Cost-per-token thinking

Model APIs quote per-million-token prices, and teams multiply by volume and call it a budget. That's the beginning of cost thinking, not the end. Realistic cost-per-token accounting includes:

  1. Input vs. output asymmetry. Output tokens cost 3–4x input tokens on most APIs because decode is the expensive phase. Prompt engineering that moves work from output to input (few-shot examples instead of long generations, structured extraction instead of free text) is a cost optimization disguised as a quality practice.
  2. The retry and cascade multiplier. If 15% of requests retry and 20% escalate to a bigger model, your effective cost per successful user request is 30–50% above the naive token math. Measure cost per completed task, not per token.
  3. Idle capacity. Self-hosted GPUs cost money at 3 AM too. If traffic has a 5x peak-to-trough ratio and you provision for peak, effective cost per token is dominated by utilization, not serving efficiency. Autoscale on tokens-per-second (not request count); use spot capacity for batch tiers.
  4. The hidden tax of context. Every token stuffed into the prompt is paid on every request and inflates the KV cache, slowing every decode step of other requests sharing the batch. Retrieval that narrows context to what's needed is usually cheaper than a bigger context window — measure it.

The meta-lesson: LLM serving cost is almost never reduced by a cleverer model choice. It's reduced by batching discipline, cache hit rates, routing, and right-sizing max tokens — the operational layer between the model and the user. Get that layer right and the same GPUs serve 3–5x the traffic.

11 North-star metrics that don't lie: picking growth metrics that survive contact with reality Goodhart's law as a gaming mechanism, input vs. output metric hierarchies, counter-metrics, and a five-check stress test for any metric. Read articleClose article

Every company has a north-star metric. Most of them are decorative — a number on a dashboard that goes up while the business quietly goes sideways. The difference between a metric that guides decisions and one that decorates them is decided before you commit to it, and it comes down to a handful of disciplines most teams skip.

Goodhart's law in practice

"When a measure becomes a target, it ceases to be a good measure." Everyone quotes it; few design against it. The mechanism is specific: the moment a metric determines bonuses, promotions, or roadmap priority, every team in the blast radius starts optimizing the measurement instead of the outcome.

I've seen the pattern enough times to recognize its shape. A marketplace makes "completed bookings" the north star; within two quarters, hosts are gaming listings and bookings get cancelled off-platform after first contact — the metric went up, the take rate went down. A SaaS company picks "weekly active users"; teams ship engagement-bait notifications and redefine "active" down to "opened the app once" — WAU climbs while revenue per user falls.

The defense isn't to avoid targets — it's to assume gaming from day one and pick metrics that are expensive to game. A metric is game-resistant when the cheapest way to move it is to do the real work: it's hard to fake retained, paying usage without delivering value. Before adopting a metric, ask the adversarial question: "If a team wanted to move this 20% without helping the business, what's the cheapest trick?"

Input vs. output metrics

Output metrics (revenue, retention, NPS) tell you whether you won. Input metrics (activation rate, time-to-first-value, support contacts per order) tell you why — and they're the ones teams can actually move week to week. The classic failure is holding teams accountable to an output metric nobody's work directly touches, then watching everyone flail.

The right structure is a hierarchy, not a single number. The north star is an output metric. Under it sit 3–5 input metrics, each owned by a team, each with a demonstrated causal link to the north star. The operative word is demonstrated: "we believe faster onboarding improves retention" is a hypothesis, not a hierarchy. You need the analysis showing that cohorts with faster time-to-value actually retain better, or you're building an org chart on vibes.

The trap on the other side is input-metric proliferation — twenty "levers" that let every team claim success while the north star flatlines. Keep the tree shallow. If an input moves and the north star doesn't follow within the expected lag, the causal link was wrong. Kill the input, not the messenger.

Leading vs. lagging indicators

Lagging indicators (churn, quarterly revenue) are accurate and useless for steering — by the time they move, the decisions that caused them are months old. Leading indicators (trial-to-paid conversion, week-1 engagement depth) move first and let you correct course, but they're noisier and easier to misread.

You need both, and you need to know which job each one does. Leading indicators run the weekly operating cadence: they're the dials. Lagging indicators run the quarterly strategy review: they're the scoreboard. Never use a lagging indicator as a dial (reorging onboarding over Q2 churn) or a leading indicator as a scoreboard (declaring victory on retention from week-1 engagement).

One discipline that pays for itself: for every leading indicator, write down its expected lag to the lagging metric it predicts, then check. Most teams never close this loop — which is how leading indicators quietly become vanity metrics with extra steps.

The failure mode of vanity metrics, with examples

A vanity metric is one that makes you feel good while telling you nothing about the health of the business. They're easy to grow, hard to act on, and disconnected from value.

  • Registered users. The graveyard metric. A million signups with 2% activation is a mailing list, not traction. The honest version is activated users — with activation defined as reaching value.
  • Total pageviews / sessions. Goes up when you add pagination, infinite scroll, or confusing navigation. A redesign that reduces pageviews by getting users to their answer faster is a win the metric calls a loss. Measure task completion, not eyeball time — unless eyeball time is literally the product.
  • Raw NPS or app-store rating. A 4.8-star app with flat revenue and rising churn is a paradox only if you trust the metric. Ratings skew toward extremes and say nothing about the silent majority quietly leaving.
  • Feature adoption counts. "10,000 users tried the new dashboard" — and 9,500 never came back. Trial without retention is curiosity. The metric that matters is sustained usage among the cohort that tried it.
  • Gross bookings / GMV without take-rate or margin. GMV can grow 40% while unit economics collapse on subsidized or low-margin transactions. Pair every scale metric with its margin counterpart.

The test is always the same: "If this metric doubled and nothing else changed, would the business be healthier?" If the answer is "not necessarily," it's vanity.

Counter-metrics: the seatbelt for your north star

Every metric you push on creates pressure somewhere else, and the pressure finds the place you're not measuring. Counter-metrics are the explicit list of things you're not willing to sacrifice to move the north star.

Pushing "checkout conversion"? Your counter-metrics are refund rate, support contacts per order, and average order margin — because the easiest way to juice conversion is discounts and dark patterns. Pushing "time spent in app"? Counter-metric: task completion rate and uninstall rate, because the easiest way to increase time spent is to make everything slower.

The discipline that makes counter-metrics work: thresholds and teeth. "We'll watch refund rates" is decoration. "If refund rate rises more than 0.5 points, the experiment rolls back regardless of conversion lift" is a guardrail.

Metric review cadence and ownership

Metrics rot. Definitions drift, the business changes, and a metric that was perfect eighteen months ago now measures a world that doesn't exist. The fix is boring and non-negotiable: a named owner for every metric in the tree, and a review cadence.

  • Weekly: input metrics, owned by the team that moves them. The question is "what did we learn and what changes this week?"
  • Monthly: the north star and its trend. The question is "is the tree still predictive?" — do input movements still show up downstream?
  • Quarterly: the definitions themselves. "Does this metric still mean what we think it means?" This is where you catch definition drift (three teams computing "active" three ways), gaming residue, and metrics the business has outgrown.

Ownership means one person whose job includes defending the metric's integrity — not a committee, not "the data team." Committees don't notice when a definition quietly changes in a dbt model. A named owner does.

How to stress-test a metric before committing

Before a metric earns a place in the tree, put it through five checks — an afternoon's work that saves quarters of misdirected effort.

  1. The gaming test. Name the three cheapest ways to move it without creating value. If any is cheap and hard to detect, add a counter-metric that catches it or pick a different metric.
  2. The definition test. Write the definition in one paragraph, including edge cases. If two reasonable people compute it two ways, it's not a metric yet — it's an argument waiting to happen.
  3. The lag test. State when you expect it to move after the underlying reality changes. A metric with unknown lag is a historian, not a dial.
  4. The cohort test. Check that it behaves on cohorts, not just in aggregate — aggregates hide Simpson's paradox, where overall conversion rises while every segment's conversion falls as the mix shifts.
  5. The "so what" test. Describe a concrete decision this metric would change. A metric with no decision attached is a dashboard ornament.

Metrics are the interface between what a company wants and what its people do all day. A good north star coordinates hundreds of small decisions toward the same outcome, because everyone can see how their work ladders up. That's worth an afternoon of stress-testing. The alternative is a number that goes up while the business goes sideways, and a team that learns to celebrate the number instead of questioning it.