Luca Barberis
AI x GTM

Tables, not prompts.
AI applied to go-to-market.

A prompt asks the model to be smart. A table tells it where to put the answer. How I run analysis, outbound, SEO, Google Ads and Meta Ads through an LLM, and why a chain of tables beats the chat window.

00

In brief

01

A table doesn't make the model smarter. It makes its mistakes visible. Ask for a plan in prose and the gaps get written around. Ask for a table and every gap is a cell that says MISSING, or a verdict you can check.

02

Five files, one chain. Analysis, outbound, SEO, Google Ads, Meta Ads. Each one is the same eight moves: route, brief, reality check, inventory, verdict, build, QA, weekly review.

03

The model is rented. The judgment is the asset, and the judgment isn't the rules. Most rules are in any decent playbook. The judgment is which evidence counts here, where the thresholds sit, and the review that changes a rule when the numbers say so.

The point

A table is a prompt with the questions fixed, the evidence required and the verdict forced. Write that, let the model fill it, and check the verdicts yourself.

01

The question

Most of the go-to-market work I do now runs through an LLM. The analysis first, then outbound, SEO, Google Ads and Meta Ads.

The tools connect in an afternoon. Getting the model to think like an operator took months.

What does an LLM need in front of it to give me a channel plan I'd put money behind?

The answer I landed on is boring. Tables.

One thing up front. I haven't run the clean test, the same brief as a prompt and as a table chain, judged on the numbers. What follows is the mechanics and the reasons, and the mechanics are what I put money through.

Not tables as output, to make the thing look tidy. Tables as the way the model thinks: one stage at a time, with a place for every answer.

The brief to myself was one line: tables applied to GTM to keep the LLM focused on what matters.

This piece is the how. Enough of it that you can build the same chain for whatever GTM job you run.

02

A prompt and a table

Start with a well-written prompt. The company, the budget, a request for a Google Ads plan with sources cited and uncertainties flagged, and ten lines of preferences.

What comes back is a document. Headings, a few keywords, a budget split that looks reasonable, a confident tone all the way through.

It reads well. That's the problem. Reading well is what the model was trained for, and a plan that reads well is not the same thing as a plan that's right.

Three things go wrong, and they go wrong quietly.

The gaps are invisible. If the model doesn't know the account's conversion volume, it writes around it. You find out when the campaign runs.

Your ten lines get diluted. Some get honoured, some quietly don't, and nothing in the document tells you which.

And you can't check it. There's no row to mark PASS or FAIL. You read 800 words and nod.

Now the same job as a table chain.

Stage one is an intake table. One row per thing the plan depends on: budget, target CPA, conversion tracking, geography, languages, brand terms.

Two columns: what we have, and what breaks if this is wrong. Every unknown is written as MISSING.

Stage two is the keyword inventory. Thirty to 150 rows, one keyword each. Intent in one column, fit in another, then a verdict in words: Keep, Negate, Park.

Stage three is the master table. Every kept keyword gets a match type, an ad group, an ad and a bid. One row, one decision. A row with no ad group is a row the model has to explain.

Then a budget split. Then a negatives list. Then a self-check that walks every rule and writes PASS or FAIL next to it.

Same model. Same information. What changed is where a mistake can hide. A gap is a cell. A skipped rule is a cell. Padding has nowhere to go.

Table 1. The same job, asked two ways. Rows are the things that differ. Columns are what happens with a normal prompt and with a table chain. This describes the mechanics, not a measured comparison.
WhatA prompt, even a good oneTable chain
Where the question livesIn your head, then in 400 words of prompt.In the column headers. The columns are the questions.
What comes backA document. Headings, a confident tone, a plan that reads well.Rows you can count and cells you can check.
Where a gap goesInto a sentence that sounds fine. You find it when the campaign runs.Into a cell that says MISSING.
Twelve constraintsSome honoured, some quietly dropped, no way to tell which.Twelve rows, a verdict in each, and you read the FAIL column.
How you check itYou read 800 words and nod.You scan the verdict column for FAIL.
When the model is wrongWrong in a paragraph you may not reread.Wrong in a cell you can point at. Still wrong.
What feeds the next stepThe whole conversation, including the model's own mistakes.One table, with its source tags, so the next stage can re-check the facts.
Two companiesReread both plans and hold them in your head.Two summaries, side by side, same 21 fields.
What you own afterA chat transcript.The rows that hold what you know. The model underneath can change.

scroll the table sideways →

A table chain is still prompting. Four things change.

The questions are fixed in the columns. Evidence is required per cell. The verdict is forced into a closed set. And the job runs in sequence, one stage at a time.

A long prompt can ask for all of that. What it can't do is show you what was left out. A cell can, because an empty cell is a visible fact and a missing paragraph isn't.

And a table doesn't stop the model being wrong. It can fill every cell with plausible nonsense.

What it can't do is hide where it did it. That's the gain: mistakes you can find in twenty seconds, not on a second read.

The words in the cells matter too. PASS and FAIL. Keep and Negate. A fixed vocabulary stops the model inventing a fourth verdict that sounds like a compromise.

And the output of one table is the input to the next. The model never carries the whole job in its head. It carries one table.

That cuts both ways. A mistake in stage two walks into stage three with a straight face. It's why the source tags travel with the row: the next stage can re-check the fact, not just inherit the verdict.

Still, one table at a time is the part I'd underline. A lot of what goes wrong on long jobs is the model losing the plot partway through.

A chain of small tables has no partway. There's this table, and then the next one.

03

The tables

Analysis runs first, always. Four questions in ten stages. What the business does and why anyone pays for it. Who buys. Who else sells to them, and how. Which channels come first.

It ends in a master table that works for every business and that can be used to summarize every business.

Twenty-one fields, the same for a retail SaaS in Zurich and a food app in Jakarta. Two companies sit side by side. A third joins as a column.

Channels get ranked last. A ranking built on fit is an opinion. Competitor evidence turns it into a bet: how many SDRs they hire, what's in their ad library, what they rank for.

Three or more SDRs at a competitor and outbound is worth a bet in that category. It says the category funds outbound, not that outbound pays back there. It's the cheapest signal I know, and it's a starting bet, not a verdict.

Every channel stays on the table for every business. One exception: no cold email to consumers.

Outbound starts from one line: the core criteria has to be RELEVANCE.

One ICP is one industry plus one title. Title is a cheap proxy for the seat on the buying committee, and the seat is what the email is written for.

A VP Sales and a CRO can share a row. A VP Sales and a Head of IT can't, and an email written for both lands with neither.

Apollo, Lusha and Hunter are simply LinkedIn big scrapers, but the data is there (who does what), paired with email guessing. So you segment on the titles, and you only send to the emails that verified.

Subject lines start from a number and a name the reader knows. That's the default I test against, not a law. Authority is numbers and names.

And a how-it-works in three or four steps beats any adjective. A mechanism reads as true; a claim reads as marketing.

SEO reads intent from the live results page, logged out and in the target country, never from the keyword.

Google Ads keeps campaigns for budget, bidding, geo and language, never as folders for themes. Ad groups do the relevance work.

Meta runs a small set of campaigns with one job each. Four today. The count follows the platform; the one-job rule doesn't. Table 2 has the rest.

Table 2. The five files. Rows are the files. Columns are the job each one does and the chain it runs, the words its verdicts use, and the one rule that matters most in it.
FileThe job and the chainVerdict wordsThe one rule
Analysis10 stagesUnderstand a company before any channel work. Ten stages, a 21-field master summary, an 18-check QA table.Now / Build / Support / Later. PASS / FAIL. MISSING.Rank channels only after competitor evidence. Fit is an opinion; a competitor's hiring is a bet.
Outbound9 stagesCold email relevant to one ICP at a time. Brief, data check, ICP matrix, proof inventory, subjects, build, six touches, 14-check pre-send QA, weekly review.PASS / FAIL. Verified / MISSING.One industry and one title per row. Nothing else gets a campaign.
SEOT0 to T9Which page, for which query, and whether it's worth it. Brief, keyword universe, intent from the live SERP, demand, winnability, verdict, page map, packaging, on-page checklist, assumptions.Go now / Quick win / Build toward / Skip.Read intent from a clean, live results page, never from the keyword's wording.
Google Ads Search7 stagesBuy the right searches at a bid, one relevant ad each. Intake, keyword inventory of 30 to 150 rows, master table, budget split, negatives, details, self-check. Plus an audit mode.Keep / Negate / Park. Brand / Core / Competitor / Discovery.Campaigns hold budget, bidding and settings, never themes. Relevance lives in the ad group.
Meta Ads6 modesFour campaigns, creative that matches the audience and is refreshed before it saturates. Build, audit, diagnose, creative plan, scale, weekly review.Test / Scale / Retarget / Retain. [M] [P] [H] [A] [?]Never invent account data. Unknown goes in the cell and into the open questions.

scroll the table sideways →

Above every table, four rules.

Say what a row is. Fill every cell, or write MISSING. Tag every fact with where it came from. Give verdicts in words.

A fact with no tag is an assumption. Prose is two lines between tables, never a paragraph that restates one.

Gaps show up as MISSING instead of hiding inside confident paragraphs. That's most of the trick.

The rest is one row at the end of every stage: open questions. A table only finds what it asks for, and the thing that isn't a column needs somewhere to land.

Chart 1. The chain: the analysis file feeds four channel files, and every stage ends in a table

Rules: Rows / Columns line · MISSING, never a guess · a source tag on every fact · PASS / FAIL words · no prose restating a table Analysis A1 to A4: what, model, personas, geography C, D: competitors and their GTM B: 20 channels scored, 4 Now max 21-field master summary OutboundICP matrix · relevance · proof · build · 6 touches · QA SEOuniverse · intent · verdict · page map · checklist Google Ads Searchkeywords · master table · budget split · negatives Meta Ads4 campaigns · saturation ladder · decisions Weekly review thresholds · actions · open questions
Takeaway: the analysis output is the input to every channel file, and every channel file ends in a weekly review table. The rules bar at the top applies to every table in every file.
04

What the research says, and what it doesn't

The mechanics above are enough to run on. The research explains the problem they answer. It doesn't prove the answer, and I'll keep the claims to what was measured.

Models don't read long inputs evenly. Put the answer in the middle of twenty documents and accuracy drops; with enough documents GPT-3.5 did worse than with none at all.[1]

Chroma tested 18 current models in 2025 and every one got worse as input grew, even on trivial tasks. Shuffling the text so it lost its flow made retrieval better.[2]

The narrow lesson is the useful one. Isolate the next question, keep the format stable, and don't put the whole job in one window. A table chain does those three things. It doesn't follow that prose is the wrong format for everything.

They're touchy about format. Change spacing and separators, same meaning, and accuracy swung by up to 76 points in one study and 40% in another.[3][4]

Fixing the format removes one source of noise. It doesn't pick the best format for you: the same study found the format that works best differs by model, so swapping the model means re-checking the tables, not just plugging it in.

Instructions get dropped as they pile up. Given 500 constraints, the best model in mid-2025 kept 68% and favoured the early ones.[5] Twelve rules in a brief is a smaller problem, and newer models do better.

The QA table is there for a different reason. A rule with its own row and its own verdict gets checked. A rule buried in a prompt gets averaged in with everything else.

And errors compound. Get each step right 90% of the time and a ten-step job fails two times in three.[6] Small tables, a verdict at the end of each, nothing unverified carried forward. Chart 2 is why.

The same paper shows reasoning models stretching that horizon on their own. One stage per table is a 2025 answer to a 2025 problem, and section 06 says which rows I expect to drop as models improve.

Table 3. Four ways a model drifts, and the table rule that answers each. Rows are the failure modes. Columns are what was measured, and the rule that exists because of it.
FailureWhat was measuredRule in the files
Loses the middleLiu 2023; Chroma 2025Accuracy drops for anything placed mid-context, and drops further as input grows. All 18 models tested; shuffled text beat coherent text.[1][2]One table per stage. The input to stage N is the table from stage N-1, not the conversation so far.
Touchy about formatSclar 2023; He 2024Same meaning, different spacing or separators: swings of up to 76 accuracy points in one study, up to 40% in another.[3][4]One fixed format, the same columns and verdict words at every stage, re-checked when the model changes.
Drops instructionsIFScale 2025Given 500 constraints, the best model kept 68% and favoured the early ones.[5]Rules live in a QA table run after the build, one row and one verdict each.
Compounds its errorsSinha 2025Success over t steps falls as p to the power t. Per-step accuracy falls as the job goes on, and models copy their own earlier mistakes.[6]A verdict closes every stage, and the source tags travel with the row so the next stage can re-check the fact.

scroll the table sideways →

Chart 2. How many steps a model gets through before it fails, as a function of single-step accuracy

0204060 steps completed 80%85%90%95%99% single-step accuracy p 69 steps at 50% success 22 steps at 80% success 6.613.5 Steps H at which success stays above s, with H = ln(s) / ln(p). At p = 90%, two ten-step tasks in three fail.
Takeaway: at 90% per-step accuracy a model gets through about 7 steps half the time; at 99% about 69. A ten-stage analysis is a long job, and small gains in step accuracy are what buy it. The chart doesn't show that tables raise step accuracy; it shows that if they raise it a little, long jobs get much easier. Curve computed from Sinha et al.[6], who assume a constant step accuracy; their own data shows it falling as the job goes on, so the real curve is worse.

Humans found the same thing years ago. A 19-item surgical checklist cut deaths from 1.5% to 0.8% across eight hospitals.[7] Structured interviews predict job performance about twice as well as unstructured ones.[8]

Same move each time. Fix the question, force the answer.

05

One skeleton, five channels

Put the five files next to each other and the same eight moves show up in every one.

Nothing new in the eight. Any careful channel runs them. The point is making the model run them in that order and stop at each one.

Route the request. Take the brief. Check the data against reality. Build the inventory. Score it and say the verdict in words. Build the deliverable. Run the QA table. Review weekly against thresholds.

The columns change with the channel. The chain doesn't.

Table 4. The same eight moves in all five files. Rows are the moves in the order they run. Columns are what each move is called, and what it produces, in each file.
The moveAnalysisOutboundSEOGoogle Ads SearchMeta Ads
Routewhich stages run for this requestStage 0, 9 request typesStage 0, 7 request typesStep 1, 4 modesBuild workflow or audit mode6 modes, each a table sequence
Intakethe brief, gaps marked9 inputs, Have / MISSING; research first, ask once10 inputs, incl. the free assets the sender can giveT0 brief, 10 fields, Confirmed / Assumed / MissingInputs with "what breaks if this is wrong"13 fields with "blocking if unknown"
Reality checkwhat the data is worth11 source tags; [Illustrative] banned in real workField-by-field trust in Apollo, Lusha, Hunter; send only to verified emailsTool metrics are estimates; Search Console is the only first-party dataRules vetted against Google's documentationEvidence tags [M] [P] [H] [A] [?]; heuristics never pass as platform rules
Inventorythe universe to choose fromA1 to A4, competitor register, D1 to D4 competitor GTMICP matrix, one industry and one title per row; proof inventoryT1 keyword universe; T2 intent from the live SERPKeyword inventory, 30 to 150 rowsAudience map; concept matrix; persona by stage grid
Verdictscore it, say it in wordsB2: Impact x Likelihood, 1 to 25; Now / Build / Support / LaterRelevance test, five PASS / FAIL columns per ICP; subject checkT5: value x winnability x demand; Go now / Quick win / Build toward / SkipKeep / Negate / Park; tier Brand / Core / Competitor / DiscoveryDecision rules: kill, graduate, scale, refresh, consolidate, cap
Buildthe deliverable21-field master summary; B4 test plan with kill criteriaEmail build table, six parts each PASS / FAIL; six touches over 21 daysT6 page map; T7 title, H1, meta, slug with character countsMaster table: keyword, match type, intent, ad group, ad text, bid; budget splitCampaign plan, creative quota, budget step schedule
QAbefore anything shipsStage 10: 18 checks; deliver only when all PASSStage 8: 14 pre-send checksStep 5: 12-line self-checkStage 7 self-check incl. character limits by scriptMeasurement check; every answer ends in Decisions and Open questions
Weekly reviewthresholds and one action eachOpen questions; update mode touches only what new information changesStage 9: bounce, reply, positive; one action per thresholdT9 assumptions and gaps; judge changes only after recrawlAudit mode: findings, corrected master table, diffKPI summary, ad-level saturation ladder, rules fired, next week's actions

scroll the table sideways →

The thresholds are operating points, not discoveries. The market data says they sit in a sane range.

Instantly's 2026 report has the average cold email reply rate at 3.43%.[9] Belkins, counting replies against total sends to strangers, gets 0.45%.[11] Two thermometers, not one trend.

Small lists beat big ones. 5.8% reply under 50 recipients, 2.1% above 500.[12]

Top senders keep bounces under 2%. The average is above 5%. Google's complaint ceiling is 0.3%.[10][12][13]

The weekly review acts at 2% bounce, 1.5% replies and 1% positives, after 300 delivered. Each threshold names the first thing to change, not the only cause. A bad list shows up in all three.

Two rows I'd add: a complaint threshold, and a shorter sequence, since four or more emails triple the complaints.[14]

Table 5. The outbound thresholds against published benchmarks. Rows are the metrics the weekly review watches. Columns are the rule as written in the file, the market number, and how the two read together.
MetricThe operating pointMarket numberRead
BouncePause above 2%. Re-verify, drop or batch catch-all domains.Top performers under 2%; the average is 5.1%.[10][12]Aligned
Reply, any kindUnder 1.5% after 300 delivered: the subject is the first thing to change, then the list.3.43% across all senders (Instantly 2026); 0.45% against total sends (Belkins 2025).[9][11]A floor to act on
Positive replyUnder 1% after 300: the CTA asset is the first suspect, then the offer, then the list.No published split. Positive share is the number that ties to pipeline.[11]No benchmark
Clone a hookAbove 3% positive: same title, adjacent industry.Instantly calls 5 to 10% total reply good, 15%+ excellent on focused plays.[10]Conservative
List sizeOne ICP per campaign. One industry, one title.Under 50 recipients 5.8% reply; above 500, 2.1%.[12]Same direction
Spam complaintsNo threshold in the file yet.Google's ceiling is 0.3%, recommended under 0.1%.[13]Row to add
Sequence lengthSix touches over 21 days across three threads; a call only from touch five.Reply gains concentrate in the first follow-ups; four or more emails more than triple unsubscribes and complaints.[14]Six vs four

scroll the table sideways →

06

Where the value sits

Three layers. The tools, the model, the files.

The tools are plumbing. Apollo, Instantly, Ahrefs, Search Console, the ad platforms, connected wherever there's an API.

The model is rented by the token. A better one arrives every few months, and it swaps in without touching a table.

The files are where the value is. But not all of the files, and it took me a while to see the split.

Some rows exist to cover for the model. One table per stage, so nothing gets lost in the middle. A QA table, because the model drops constraints. Verdict words, because it invents compromises.

Every release needs fewer of those rows.

Some stay anyway, because I need them. To compare two plans on the same rows. To approve a plan in a minute. To see later why a decision was made.

Other rows carry judgment. Not the rules themselves; most of those you can read in any decent playbook, and a good model has read the playbooks.

The judgment is which evidence counts for this business. Where the thresholds sit. When a rule has stopped working, which only the weekly review can tell you.

So the asset isn't the tables, and it isn't the rules either. It's the loop: rules, numbers, revised rules. The files are where the loop lives today, and the format will change before the loop does.

The format does help, for now. A skill costs about 100 tokens until it's needed, then a file under 5,000, then references only when a stage asks for them.[15] Karpathy called the job "filling the context window with just the right information for the next step".[16]

Outbound is where the loop matters most. AI SDRs send more emails from the same prompt. Nothing in the prompt reads the reply column and changes the next one.

Table 6. What the rows are for. Rows are rules taken from the five files. Columns are who the rule serves, why it exists, and what happens to it as models improve. Rows for the model cover its weaknesses; rows for the reader let me check and compare; judgment rows carry a call and the number that tests it.
RuleServesWhy it existsWhen a better model arrives
One table per stageFor the modelModels lose the middle of long inputs.Goes, once the model holds a full job without drifting.
QA table after the buildFor the model For the readerEach rule gets its own row and its own verdict.Shrinks for the model. Stays as the audit trail I read.
Verdicts in wordsFor the model For the readerStops invented compromises; keeps outputs comparable.Stays. Two companies on the same verdict words is how I compare them.
MISSING and source tagsFor the model JudgmentThe model won't flag a gap unless told to. And I want to see it.Stays. No model is trained to prefer admitting a gap over filling it.
Open-questions row at the end of every stageJudgmentA table finds only what it asks for.Stays. The better the model, the more it can put there.
Competitor evidence before channel rankingJudgmentFit is an opinion. A competitor's hiring is a bet.Stays. A better model still doesn't know what counts as evidence.
Three or more SDRs means outbound is worth a betJudgmentThe category funds the motion. Cheapest signal there is; not proof it pays back.Stays as the starting bet. The weekly review decides.
One industry, one title per rowJudgmentTitle is the proxy for the buying-committee seat the email is written for.Stays as the default. Merge rows when two titles hold the same seat.
Intent from the live SERP, not the keywordJudgmentGoogle decides what a query means, not the wording. Logged out, target country.Stays.
Campaigns hold budget and settings, never themesJudgmentAd groups do the relevance work. Campaigns hold budget, bidding, geo, language.Stays until Google moves the controls.
Thresholds in the weekly reviewJudgmentOperating points from my own numbers, each naming the first thing to change.Change whenever the numbers say so. That's what they're for.

scroll the table sideways →

07

Where this breaks

Four objections worth taking seriously.

Sutton's Bitter Lesson. In AI, methods that lean on computation keep beating methods that encode human knowledge.[17] Half of these files is human knowledge. The other half exists to cover for the model's weaknesses.

If the next model reads a company page and ranks twenty channels on its own, those rows go, and they should. The judgment rows stay until the market changes.

Forcing a strict output format can make a model reason worse. Tam and colleagues measured it: the tighter the format, the worse the reasoning, and strict JSON was the worst case.[18]

So the files separate the two steps. The model explains its reasoning in prose first, then it fills the verdict cell. The format constrains the answer, not the thinking behind it.

A table that isn't used changes nothing. When Ontario made the surgical checklist mandatory in 101 hospitals, mortality didn't move. Follow-up work found the checklist was mostly not being run in the operating room.[19][20]

The same applies here. The chain only works if every job goes through it, with no shortcut around it.

The verdict column can be wrong. A QA table only finds the gaps it was told to look for.

Someone with the sources in front of them has to check the rest. The tags on each row are there so that takes minutes, not an afternoon.

08

Takeaways

  • A table is a prompt with the questions fixed, the evidence required and the verdict forced. Write that.
  • Check the verdicts. A table makes mistakes visible; it doesn't make them go away.
  • Four rules over every table: say what a row is, fill every cell or write MISSING, tag every fact, give verdicts in words.
  • Same skeleton for every channel: route, brief, reality check, inventory, verdict, build, QA, weekly review.
  • Analysis before channels. Competitors before ranking. Every channel stays on the table except cold email to consumers.
  • Outbound is relevance. One buying-committee seat per row, with industry and title as the proxy; a number and a name in the subject; a how-it-works in three or four steps; verified emails only.
  • Every stage ends in an open-questions row. The thing that isn't a column needs somewhere to land.
  • The model is rented. The asset is the loop: rules, numbers, revised rules. Rows that cover for the model go when it stops needing them. Rows for the reader and judgment rows stay.

Some rows are there for the model, and they'll go. Some are there for me, so I can read and approve the work. The rest carry a rule and the number that tests it. Those are the ones worth writing down.

Notes & sources

Stage counts, thresholds and rules are described from the five files as they stand in September 2026. Everything else traces to the sources below, numbered in order of first appearance.

  1. Nelson F. Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2023, arXiv 2307.03172. arxiv.org/abs/2307.03172
  2. Kelly Hong, Anton Troynikov and Jeff Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance", Chroma, 14 July 2025. trychroma.com/research/context-rot
  3. Melanie Sclar, Yejin Choi, Yulia Tsvetkov and Alane Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design", ICLR 2024, arXiv 2310.11324. arxiv.org/abs/2310.11324
  4. Jia He et al., "Does Prompt Formatting Have Any Impact on LLM Performance?", arXiv 2411.10541, November 2024. arxiv.org/abs/2411.10541
  5. Daniel Jaroslawicz et al., "How Many Instructions Can LLMs Follow at Once?" (IFScale), arXiv 2507.11538, 2025. arxiv.org/abs/2507.11538
  6. Akshit Sinha et al., "The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs", ICLR 2026, arXiv 2509.09677. arxiv.org/abs/2509.09677
  7. Alex B. Haynes et al., "A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population", New England Journal of Medicine 360(5), 2009, 491-499. doi.org/10.1056/NEJMsa0810119
  8. Paul R. Sackett, Charlene Zhang, Christopher M. Berry and Filip Lievens, "Revisiting Meta-Analytic Estimates of Validity in Personnel Selection", Journal of Applied Psychology 107(11), 2022 (structured interviews .42, unstructured .19). doi.org/10.1037/apl0000994
  9. Reachoutly, "Cold Email Response Rate (2026 Guide)", May 2026, summarising Instantly's 2026 benchmark report. reachoutly.com/cold-email/response-rate/
  10. Instantly, "Cold Email Response Rates: B2B Benchmarks", July 2026. instantly.ai/blog/cold-email-reply-rate-benchmarks/
  11. Belkins, "What are B2B Cold Email Response Rates? (2026 Study)", June 2026. belkins.io/blog/cold-email-response-rates
  12. Woodpecker, "Cold Email Statistics Based on Sending Over 20M Cold Emails", June 2026. woodpecker.co/blog/cold-email-statistics/
  13. Harbor BD, "AI SDR Results 2026: Why Automated Outbound Is Failing B2B Pipeline", April 2026. harborbd.com/blogs/ai-sdr-outbound-results-2026
  14. Growth Engineer, "Cold Email Reply Rate Benchmarks 2026: Data From 14M Sent Emails", May 2026. growthengineer.ai/blog/cold-email-reply-rate-benchmarks-2026
  15. MemoryPlugin, "Agent skills, explained: the folder format every AI tool now reads", 2026. blog.memoryplugin.com/what-are-agent-skills/
  16. Andrej Karpathy on X, 25 June 2025, as quoted in Mr Prompts, "Context Engineering". mrprompts.substack.com/p/context-engineering
  17. Rich Sutton, "The Bitter Lesson", 13 March 2019. mlanthology.org/misc/2019/sutton2019misc-bitter
  18. Zhi Rui Tam et al., "Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance", EMNLP 2024 Industry Track. aclanthology.org/2024.emnlp-industry.91
  19. David R. Urbach et al., "Introduction of Surgical Safety Checklists in Ontario, Canada", New England Journal of Medicine 370(11), 2014, as summarised by AHRQ PSNet. psnet.ahrq.gov/issue/introduction-surgical-safety-checklists-ontario-canada
  20. Lucian L. Leape, editorial accompanying Urbach et al., NEJM 2014, as reported in The Hospitalist. blogs.the-hospitalist.org/content/surgical-checklists-failed-improve-outcomes