Tables, not prompts.
AI applied to go-to-market.
A prompt asks the model to be smart. A table tells it where to put the answer. How I run analysis, outbound, SEO, Google Ads and Meta Ads through an LLM, and why a chain of tables beats the chat window.
In brief
A table doesn't make the model smarter. It makes its mistakes visible. Ask for a plan in prose and the gaps get written around. Ask for a table and every gap is a cell that says MISSING, or a verdict you can check.
Five files, one chain. Analysis, outbound, SEO, Google Ads, Meta Ads. Each one is the same eight moves: route, brief, reality check, inventory, verdict, build, QA, weekly review.
The model is rented. The judgment is the asset, and the judgment isn't the rules. Most rules are in any decent playbook. The judgment is which evidence counts here, where the thresholds sit, and the review that changes a rule when the numbers say so.
A table is a prompt with the questions fixed, the evidence required and the verdict forced. Write that, let the model fill it, and check the verdicts yourself.
The question
Most of the go-to-market work I do now runs through an LLM. The analysis first, then outbound, SEO, Google Ads and Meta Ads.
The tools connect in an afternoon. Getting the model to think like an operator took months.
What does an LLM need in front of it to give me a channel plan I'd put money behind?
The answer I landed on is boring. Tables.
One thing up front. I haven't run the clean test, the same brief as a prompt and as a table chain, judged on the numbers. What follows is the mechanics and the reasons, and the mechanics are what I put money through.
Not tables as output, to make the thing look tidy. Tables as the way the model thinks: one stage at a time, with a place for every answer.
The brief to myself was one line: tables applied to GTM to keep the LLM focused on what matters.
This piece is the how. Enough of it that you can build the same chain for whatever GTM job you run.
A prompt and a table
Start with a well-written prompt. The company, the budget, a request for a Google Ads plan with sources cited and uncertainties flagged, and ten lines of preferences.
What comes back is a document. Headings, a few keywords, a budget split that looks reasonable, a confident tone all the way through.
It reads well. That's the problem. Reading well is what the model was trained for, and a plan that reads well is not the same thing as a plan that's right.
Three things go wrong, and they go wrong quietly.
The gaps are invisible. If the model doesn't know the account's conversion volume, it writes around it. You find out when the campaign runs.
Your ten lines get diluted. Some get honoured, some quietly don't, and nothing in the document tells you which.
And you can't check it. There's no row to mark PASS or FAIL. You read 800 words and nod.
Now the same job as a table chain.
Stage one is an intake table. One row per thing the plan depends on: budget, target CPA, conversion tracking, geography, languages, brand terms.
Two columns: what we have, and what breaks if this is wrong. Every unknown is written as MISSING.
Stage two is the keyword inventory. Thirty to 150 rows, one keyword each. Intent in one column, fit in another, then a verdict in words: Keep, Negate, Park.
Stage three is the master table. Every kept keyword gets a match type, an ad group, an ad and a bid. One row, one decision. A row with no ad group is a row the model has to explain.
Then a budget split. Then a negatives list. Then a self-check that walks every rule and writes PASS or FAIL next to it.
Same model. Same information. What changed is where a mistake can hide. A gap is a cell. A skipped rule is a cell. Padding has nowhere to go.
| What | A prompt, even a good one | Table chain |
|---|---|---|
| Where the question lives | In your head, then in 400 words of prompt. | In the column headers. The columns are the questions. |
| What comes back | A document. Headings, a confident tone, a plan that reads well. | Rows you can count and cells you can check. |
| Where a gap goes | Into a sentence that sounds fine. You find it when the campaign runs. | Into a cell that says MISSING. |
| Twelve constraints | Some honoured, some quietly dropped, no way to tell which. | Twelve rows, a verdict in each, and you read the FAIL column. |
| How you check it | You read 800 words and nod. | You scan the verdict column for FAIL. |
| When the model is wrong | Wrong in a paragraph you may not reread. | Wrong in a cell you can point at. Still wrong. |
| What feeds the next step | The whole conversation, including the model's own mistakes. | One table, with its source tags, so the next stage can re-check the facts. |
| Two companies | Reread both plans and hold them in your head. | Two summaries, side by side, same 21 fields. |
| What you own after | A chat transcript. | The rows that hold what you know. The model underneath can change. |
scroll the table sideways →
A table chain is still prompting. Four things change.
The questions are fixed in the columns. Evidence is required per cell. The verdict is forced into a closed set. And the job runs in sequence, one stage at a time.
A long prompt can ask for all of that. What it can't do is show you what was left out. A cell can, because an empty cell is a visible fact and a missing paragraph isn't.
And a table doesn't stop the model being wrong. It can fill every cell with plausible nonsense.
What it can't do is hide where it did it. That's the gain: mistakes you can find in twenty seconds, not on a second read.
The words in the cells matter too. PASS and FAIL. Keep and Negate. A fixed vocabulary stops the model inventing a fourth verdict that sounds like a compromise.
And the output of one table is the input to the next. The model never carries the whole job in its head. It carries one table.
That cuts both ways. A mistake in stage two walks into stage three with a straight face. It's why the source tags travel with the row: the next stage can re-check the fact, not just inherit the verdict.
Still, one table at a time is the part I'd underline. A lot of what goes wrong on long jobs is the model losing the plot partway through.
A chain of small tables has no partway. There's this table, and then the next one.
The tables
Analysis runs first, always. Four questions in ten stages. What the business does and why anyone pays for it. Who buys. Who else sells to them, and how. Which channels come first.
It ends in a master table that works for every business and that can be used to summarize every business.
Twenty-one fields, the same for a retail SaaS in Zurich and a food app in Jakarta. Two companies sit side by side. A third joins as a column.
Channels get ranked last. A ranking built on fit is an opinion. Competitor evidence turns it into a bet: how many SDRs they hire, what's in their ad library, what they rank for.
Three or more SDRs at a competitor and outbound is worth a bet in that category. It says the category funds outbound, not that outbound pays back there. It's the cheapest signal I know, and it's a starting bet, not a verdict.
Every channel stays on the table for every business. One exception: no cold email to consumers.
Outbound starts from one line: the core criteria has to be RELEVANCE.
One ICP is one industry plus one title. Title is a cheap proxy for the seat on the buying committee, and the seat is what the email is written for.
A VP Sales and a CRO can share a row. A VP Sales and a Head of IT can't, and an email written for both lands with neither.
Apollo, Lusha and Hunter are simply LinkedIn big scrapers, but the data is there (who does what), paired with email guessing. So you segment on the titles, and you only send to the emails that verified.
Subject lines start from a number and a name the reader knows. That's the default I test against, not a law. Authority is numbers and names.
And a how-it-works in three or four steps beats any adjective. A mechanism reads as true; a claim reads as marketing.
SEO reads intent from the live results page, logged out and in the target country, never from the keyword.
Google Ads keeps campaigns for budget, bidding, geo and language, never as folders for themes. Ad groups do the relevance work.
Meta runs a small set of campaigns with one job each. Four today. The count follows the platform; the one-job rule doesn't. Table 2 has the rest.
| File | The job and the chain | Verdict words | The one rule |
|---|---|---|---|
| Analysis | Understand a company before any channel work. Ten stages, a 21-field master summary, an 18-check QA table. | Now / Build / Support / Later. PASS / FAIL. MISSING. | Rank channels only after competitor evidence. Fit is an opinion; a competitor's hiring is a bet. |
| Outbound | Cold email relevant to one ICP at a time. Brief, data check, ICP matrix, proof inventory, subjects, build, six touches, 14-check pre-send QA, weekly review. | PASS / FAIL. Verified / MISSING. | One industry and one title per row. Nothing else gets a campaign. |
| SEO | Which page, for which query, and whether it's worth it. Brief, keyword universe, intent from the live SERP, demand, winnability, verdict, page map, packaging, on-page checklist, assumptions. | Go now / Quick win / Build toward / Skip. | Read intent from a clean, live results page, never from the keyword's wording. |
| Google Ads Search | Buy the right searches at a bid, one relevant ad each. Intake, keyword inventory of 30 to 150 rows, master table, budget split, negatives, details, self-check. Plus an audit mode. | Keep / Negate / Park. Brand / Core / Competitor / Discovery. | Campaigns hold budget, bidding and settings, never themes. Relevance lives in the ad group. |
| Meta Ads | Four campaigns, creative that matches the audience and is refreshed before it saturates. Build, audit, diagnose, creative plan, scale, weekly review. | Test / Scale / Retarget / Retain. [M] [P] [H] [A] [?] | Never invent account data. Unknown goes in the cell and into the open questions. |
scroll the table sideways →
Above every table, four rules.
Say what a row is. Fill every cell, or write MISSING. Tag every fact with where it came from. Give verdicts in words.
A fact with no tag is an assumption. Prose is two lines between tables, never a paragraph that restates one.
Gaps show up as MISSING instead of hiding inside confident paragraphs. That's most of the trick.
The rest is one row at the end of every stage: open questions. A table only finds what it asks for, and the thing that isn't a column needs somewhere to land.
Chart 1. The chain: the analysis file feeds four channel files, and every stage ends in a table
What the research says, and what it doesn't
The mechanics above are enough to run on. The research explains the problem they answer. It doesn't prove the answer, and I'll keep the claims to what was measured.
Models don't read long inputs evenly. Put the answer in the middle of twenty documents and accuracy drops; with enough documents GPT-3.5 did worse than with none at all.[1]
Chroma tested 18 current models in 2025 and every one got worse as input grew, even on trivial tasks. Shuffling the text so it lost its flow made retrieval better.[2]
The narrow lesson is the useful one. Isolate the next question, keep the format stable, and don't put the whole job in one window. A table chain does those three things. It doesn't follow that prose is the wrong format for everything.
They're touchy about format. Change spacing and separators, same meaning, and accuracy swung by up to 76 points in one study and 40% in another.[3][4]
Fixing the format removes one source of noise. It doesn't pick the best format for you: the same study found the format that works best differs by model, so swapping the model means re-checking the tables, not just plugging it in.
Instructions get dropped as they pile up. Given 500 constraints, the best model in mid-2025 kept 68% and favoured the early ones.[5] Twelve rules in a brief is a smaller problem, and newer models do better.
The QA table is there for a different reason. A rule with its own row and its own verdict gets checked. A rule buried in a prompt gets averaged in with everything else.
And errors compound. Get each step right 90% of the time and a ten-step job fails two times in three.[6] Small tables, a verdict at the end of each, nothing unverified carried forward. Chart 2 is why.
The same paper shows reasoning models stretching that horizon on their own. One stage per table is a 2025 answer to a 2025 problem, and section 06 says which rows I expect to drop as models improve.
| Failure | What was measured | Rule in the files |
|---|---|---|
| Loses the middle | Accuracy drops for anything placed mid-context, and drops further as input grows. All 18 models tested; shuffled text beat coherent text.[1][2] | One table per stage. The input to stage N is the table from stage N-1, not the conversation so far. |
| Touchy about format | Same meaning, different spacing or separators: swings of up to 76 accuracy points in one study, up to 40% in another.[3][4] | One fixed format, the same columns and verdict words at every stage, re-checked when the model changes. |
| Drops instructions | Given 500 constraints, the best model kept 68% and favoured the early ones.[5] | Rules live in a QA table run after the build, one row and one verdict each. |
| Compounds its errors | Success over t steps falls as p to the power t. Per-step accuracy falls as the job goes on, and models copy their own earlier mistakes.[6] | A verdict closes every stage, and the source tags travel with the row so the next stage can re-check the fact. |
scroll the table sideways →
Chart 2. How many steps a model gets through before it fails, as a function of single-step accuracy
Humans found the same thing years ago. A 19-item surgical checklist cut deaths from 1.5% to 0.8% across eight hospitals.[7] Structured interviews predict job performance about twice as well as unstructured ones.[8]
Same move each time. Fix the question, force the answer.
One skeleton, five channels
Put the five files next to each other and the same eight moves show up in every one.
Nothing new in the eight. Any careful channel runs them. The point is making the model run them in that order and stop at each one.
Route the request. Take the brief. Check the data against reality. Build the inventory. Score it and say the verdict in words. Build the deliverable. Run the QA table. Review weekly against thresholds.
The columns change with the channel. The chain doesn't.
| The move | Analysis | Outbound | SEO | Google Ads Search | Meta Ads |
|---|---|---|---|---|---|
| Routewhich stages run for this request | Stage 0, 9 request types | Stage 0, 7 request types | Step 1, 4 modes | Build workflow or audit mode | 6 modes, each a table sequence |
| Intakethe brief, gaps marked | 9 inputs, Have / MISSING; research first, ask once | 10 inputs, incl. the free assets the sender can give | T0 brief, 10 fields, Confirmed / Assumed / Missing | Inputs with "what breaks if this is wrong" | 13 fields with "blocking if unknown" |
| Reality checkwhat the data is worth | 11 source tags; [Illustrative] banned in real work | Field-by-field trust in Apollo, Lusha, Hunter; send only to verified emails | Tool metrics are estimates; Search Console is the only first-party data | Rules vetted against Google's documentation | Evidence tags [M] [P] [H] [A] [?]; heuristics never pass as platform rules |
| Inventorythe universe to choose from | A1 to A4, competitor register, D1 to D4 competitor GTM | ICP matrix, one industry and one title per row; proof inventory | T1 keyword universe; T2 intent from the live SERP | Keyword inventory, 30 to 150 rows | Audience map; concept matrix; persona by stage grid |
| Verdictscore it, say it in words | B2: Impact x Likelihood, 1 to 25; Now / Build / Support / Later | Relevance test, five PASS / FAIL columns per ICP; subject check | T5: value x winnability x demand; Go now / Quick win / Build toward / Skip | Keep / Negate / Park; tier Brand / Core / Competitor / Discovery | Decision rules: kill, graduate, scale, refresh, consolidate, cap |
| Buildthe deliverable | 21-field master summary; B4 test plan with kill criteria | Email build table, six parts each PASS / FAIL; six touches over 21 days | T6 page map; T7 title, H1, meta, slug with character counts | Master table: keyword, match type, intent, ad group, ad text, bid; budget split | Campaign plan, creative quota, budget step schedule |
| QAbefore anything ships | Stage 10: 18 checks; deliver only when all PASS | Stage 8: 14 pre-send checks | Step 5: 12-line self-check | Stage 7 self-check incl. character limits by script | Measurement check; every answer ends in Decisions and Open questions |
| Weekly reviewthresholds and one action each | Open questions; update mode touches only what new information changes | Stage 9: bounce, reply, positive; one action per threshold | T9 assumptions and gaps; judge changes only after recrawl | Audit mode: findings, corrected master table, diff | KPI summary, ad-level saturation ladder, rules fired, next week's actions |
scroll the table sideways →
The thresholds are operating points, not discoveries. The market data says they sit in a sane range.
Instantly's 2026 report has the average cold email reply rate at 3.43%.[9] Belkins, counting replies against total sends to strangers, gets 0.45%.[11] Two thermometers, not one trend.
Small lists beat big ones. 5.8% reply under 50 recipients, 2.1% above 500.[12]
Top senders keep bounces under 2%. The average is above 5%. Google's complaint ceiling is 0.3%.[10][12][13]
The weekly review acts at 2% bounce, 1.5% replies and 1% positives, after 300 delivered. Each threshold names the first thing to change, not the only cause. A bad list shows up in all three.
Two rows I'd add: a complaint threshold, and a shorter sequence, since four or more emails triple the complaints.[14]
| Metric | The operating point | Market number | Read |
|---|---|---|---|
| Bounce | Pause above 2%. Re-verify, drop or batch catch-all domains. | Top performers under 2%; the average is 5.1%.[10][12] | Aligned |
| Reply, any kind | Under 1.5% after 300 delivered: the subject is the first thing to change, then the list. | 3.43% across all senders (Instantly 2026); 0.45% against total sends (Belkins 2025).[9][11] | A floor to act on |
| Positive reply | Under 1% after 300: the CTA asset is the first suspect, then the offer, then the list. | No published split. Positive share is the number that ties to pipeline.[11] | No benchmark |
| Clone a hook | Above 3% positive: same title, adjacent industry. | Instantly calls 5 to 10% total reply good, 15%+ excellent on focused plays.[10] | Conservative |
| List size | One ICP per campaign. One industry, one title. | Under 50 recipients 5.8% reply; above 500, 2.1%.[12] | Same direction |
| Spam complaints | No threshold in the file yet. | Google's ceiling is 0.3%, recommended under 0.1%.[13] | Row to add |
| Sequence length | Six touches over 21 days across three threads; a call only from touch five. | Reply gains concentrate in the first follow-ups; four or more emails more than triple unsubscribes and complaints.[14] | Six vs four |
scroll the table sideways →
Where the value sits
Three layers. The tools, the model, the files.
The tools are plumbing. Apollo, Instantly, Ahrefs, Search Console, the ad platforms, connected wherever there's an API.
The model is rented by the token. A better one arrives every few months, and it swaps in without touching a table.
The files are where the value is. But not all of the files, and it took me a while to see the split.
Some rows exist to cover for the model. One table per stage, so nothing gets lost in the middle. A QA table, because the model drops constraints. Verdict words, because it invents compromises.
Every release needs fewer of those rows.
Some stay anyway, because I need them. To compare two plans on the same rows. To approve a plan in a minute. To see later why a decision was made.
Other rows carry judgment. Not the rules themselves; most of those you can read in any decent playbook, and a good model has read the playbooks.
The judgment is which evidence counts for this business. Where the thresholds sit. When a rule has stopped working, which only the weekly review can tell you.
So the asset isn't the tables, and it isn't the rules either. It's the loop: rules, numbers, revised rules. The files are where the loop lives today, and the format will change before the loop does.
The format does help, for now. A skill costs about 100 tokens until it's needed, then a file under 5,000, then references only when a stage asks for them.[15] Karpathy called the job "filling the context window with just the right information for the next step".[16]
Outbound is where the loop matters most. AI SDRs send more emails from the same prompt. Nothing in the prompt reads the reply column and changes the next one.
| Rule | Serves | Why it exists | When a better model arrives |
|---|---|---|---|
| One table per stage | For the model | Models lose the middle of long inputs. | Goes, once the model holds a full job without drifting. |
| QA table after the build | For the model For the reader | Each rule gets its own row and its own verdict. | Shrinks for the model. Stays as the audit trail I read. |
| Verdicts in words | For the model For the reader | Stops invented compromises; keeps outputs comparable. | Stays. Two companies on the same verdict words is how I compare them. |
| MISSING and source tags | For the model Judgment | The model won't flag a gap unless told to. And I want to see it. | Stays. No model is trained to prefer admitting a gap over filling it. |
| Open-questions row at the end of every stage | Judgment | A table finds only what it asks for. | Stays. The better the model, the more it can put there. |
| Competitor evidence before channel ranking | Judgment | Fit is an opinion. A competitor's hiring is a bet. | Stays. A better model still doesn't know what counts as evidence. |
| Three or more SDRs means outbound is worth a bet | Judgment | The category funds the motion. Cheapest signal there is; not proof it pays back. | Stays as the starting bet. The weekly review decides. |
| One industry, one title per row | Judgment | Title is the proxy for the buying-committee seat the email is written for. | Stays as the default. Merge rows when two titles hold the same seat. |
| Intent from the live SERP, not the keyword | Judgment | Google decides what a query means, not the wording. Logged out, target country. | Stays. |
| Campaigns hold budget and settings, never themes | Judgment | Ad groups do the relevance work. Campaigns hold budget, bidding, geo, language. | Stays until Google moves the controls. |
| Thresholds in the weekly review | Judgment | Operating points from my own numbers, each naming the first thing to change. | Change whenever the numbers say so. That's what they're for. |
scroll the table sideways →
Where this breaks
Four objections worth taking seriously.
Sutton's Bitter Lesson. In AI, methods that lean on computation keep beating methods that encode human knowledge.[17] Half of these files is human knowledge. The other half exists to cover for the model's weaknesses.
If the next model reads a company page and ranks twenty channels on its own, those rows go, and they should. The judgment rows stay until the market changes.
Forcing a strict output format can make a model reason worse. Tam and colleagues measured it: the tighter the format, the worse the reasoning, and strict JSON was the worst case.[18]
So the files separate the two steps. The model explains its reasoning in prose first, then it fills the verdict cell. The format constrains the answer, not the thinking behind it.
A table that isn't used changes nothing. When Ontario made the surgical checklist mandatory in 101 hospitals, mortality didn't move. Follow-up work found the checklist was mostly not being run in the operating room.[19][20]
The same applies here. The chain only works if every job goes through it, with no shortcut around it.
The verdict column can be wrong. A QA table only finds the gaps it was told to look for.
Someone with the sources in front of them has to check the rest. The tags on each row are there so that takes minutes, not an afternoon.
Takeaways
- A table is a prompt with the questions fixed, the evidence required and the verdict forced. Write that.
- Check the verdicts. A table makes mistakes visible; it doesn't make them go away.
- Four rules over every table: say what a row is, fill every cell or write MISSING, tag every fact, give verdicts in words.
- Same skeleton for every channel: route, brief, reality check, inventory, verdict, build, QA, weekly review.
- Analysis before channels. Competitors before ranking. Every channel stays on the table except cold email to consumers.
- Outbound is relevance. One buying-committee seat per row, with industry and title as the proxy; a number and a name in the subject; a how-it-works in three or four steps; verified emails only.
- Every stage ends in an open-questions row. The thing that isn't a column needs somewhere to land.
- The model is rented. The asset is the loop: rules, numbers, revised rules. Rows that cover for the model go when it stops needing them. Rows for the reader and judgment rows stay.
Some rows are there for the model, and they'll go. Some are there for me, so I can read and approve the work. The rest carry a rule and the number that tests it. Those are the ones worth writing down.
Notes & sources
Stage counts, thresholds and rules are described from the five files as they stand in September 2026. Everything else traces to the sources below, numbered in order of first appearance.
- Nelson F. Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2023, arXiv 2307.03172. arxiv.org/abs/2307.03172
- Kelly Hong, Anton Troynikov and Jeff Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance", Chroma, 14 July 2025. trychroma.com/research/context-rot
- Melanie Sclar, Yejin Choi, Yulia Tsvetkov and Alane Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design", ICLR 2024, arXiv 2310.11324. arxiv.org/abs/2310.11324
- Jia He et al., "Does Prompt Formatting Have Any Impact on LLM Performance?", arXiv 2411.10541, November 2024. arxiv.org/abs/2411.10541
- Daniel Jaroslawicz et al., "How Many Instructions Can LLMs Follow at Once?" (IFScale), arXiv 2507.11538, 2025. arxiv.org/abs/2507.11538
- Akshit Sinha et al., "The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs", ICLR 2026, arXiv 2509.09677. arxiv.org/abs/2509.09677
- Alex B. Haynes et al., "A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population", New England Journal of Medicine 360(5), 2009, 491-499. doi.org/10.1056/NEJMsa0810119
- Paul R. Sackett, Charlene Zhang, Christopher M. Berry and Filip Lievens, "Revisiting Meta-Analytic Estimates of Validity in Personnel Selection", Journal of Applied Psychology 107(11), 2022 (structured interviews .42, unstructured .19). doi.org/10.1037/apl0000994
- Reachoutly, "Cold Email Response Rate (2026 Guide)", May 2026, summarising Instantly's 2026 benchmark report. reachoutly.com/cold-email/response-rate/
- Instantly, "Cold Email Response Rates: B2B Benchmarks", July 2026. instantly.ai/blog/cold-email-reply-rate-benchmarks/
- Belkins, "What are B2B Cold Email Response Rates? (2026 Study)", June 2026. belkins.io/blog/cold-email-response-rates
- Woodpecker, "Cold Email Statistics Based on Sending Over 20M Cold Emails", June 2026. woodpecker.co/blog/cold-email-statistics/
- Harbor BD, "AI SDR Results 2026: Why Automated Outbound Is Failing B2B Pipeline", April 2026. harborbd.com/blogs/ai-sdr-outbound-results-2026
- Growth Engineer, "Cold Email Reply Rate Benchmarks 2026: Data From 14M Sent Emails", May 2026. growthengineer.ai/blog/cold-email-reply-rate-benchmarks-2026
- MemoryPlugin, "Agent skills, explained: the folder format every AI tool now reads", 2026. blog.memoryplugin.com/what-are-agent-skills/
- Andrej Karpathy on X, 25 June 2025, as quoted in Mr Prompts, "Context Engineering". mrprompts.substack.com/p/context-engineering
- Rich Sutton, "The Bitter Lesson", 13 March 2019. mlanthology.org/misc/2019/sutton2019misc-bitter
- Zhi Rui Tam et al., "Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance", EMNLP 2024 Industry Track. aclanthology.org/2024.emnlp-industry.91
- David R. Urbach et al., "Introduction of Surgical Safety Checklists in Ontario, Canada", New England Journal of Medicine 370(11), 2014, as summarised by AHRQ PSNet. psnet.ahrq.gov/issue/introduction-surgical-safety-checklists-ontario-canada
- Lucian L. Leape, editorial accompanying Urbach et al., NEJM 2014, as reported in The Hospitalist. blogs.the-hospitalist.org/content/surgical-checklists-failed-improve-outcomes