LAB 00 · EDGE● READY

Where it breaks

OUTCOME / EVERY FAILURE SHAPE, RECOGNIZED

STACK
ChatGPT · Claude
LEVEL
Beginner
TIME
35 minutes
PLATFORMWhich product’s names this page uses — BOTH shows every name.

PULLS OFF THE SHELFThe tool call Context rot Steering the average Giving absence a word

What you'll walk away with: the six failure shapes — each recognized from your chair, each answered by a window move — and the checklist every lab after this one runs on.

You need: nothing today — no account; both demos run in your browser. About 35 minutes.

Before the edge, thirty seconds of retrieval — from memory, not from the page: what is the machine, and what feeds it? Say it, then check. The machine: odds over a window, rebuilt on every send. The workspace: fixtures that all load that same window. If either sentence wouldn't come, that lesson is worth the re-visit before this one. This lesson walks the edge — where the machine's one move stops working, what each failure looks like from your chair, and the move that answers it.

One thing this lesson is not, because it cannot honestly be: a list of what AI is good at and bad at. Any such list is dated the day it is written, wrong within the year — and the boundary between tasks these machines handle and tasks they fumble cannot be deduced in advance, even by the people who build them. You map that edge by use, on your own material, exactly as the first lesson said. What can be learned in advance is the failure shapes. They fall straight out of the mechanism you watched — the odds, the window, the pile — which is why they show up in every product, under every brand, and have survived every upgrade so far. Learn the shapes and mapping-by-use gets cheap: you recognize the terrain the moment you cross the line.

No account is needed here; both demos run in your browser. About 35 minutes.

Two ways to an answer

Start with arithmetic, because nowhere is the mechanism easier to catch bare-handed.

Ask for 2 + 2 and "4" dominates the odds the way "you" follows "thank": that exact continuation sits in the training pile beyond counting. Nothing you can audit ran — no visible procedure, no receipt; the digits of an answer are picked the way every other token is picked, by the odds over what usually comes next.

So predict, before you touch the demo: same machine, but the question is 37,481 × 2,919 — a product that has almost certainly never appeared in any text anywhere. What do the odds do?

TOY — LOADING…

What you watched: the odds still produce something answer-shaped — right length, right leading digits, even the right final digit, because local patterns are exactly what odds are good at. The true product is in the odds too, and it is not the favorite. And the part that matters more than any wrong number: a computed result and a fluent guess look typographically identical. Nothing in the answer — not the confidence, not the formatting — tells you which one you got.

Now flip the demo to the tool path and watch the same request go a different way. The model writes a request — more text, in the same window. The app runs it as real code. The true product comes back into the window, and the reply continues from it. You met this door once before, when web search filled the window in the first lesson; it is the same door, and the tool call is its general name:

A TOOL CALL — REACHING PAST THE ODDSSCHEMATIC — THE SHAPE, NOT THE SIZES
THE MODELwrites a request — more text in the window
THE APPactually runs it — search, code, files
THE RESULTpasted back into the window as text
THE REPLYcontinues from what came back

No visible tool step? Treat it as unrun — “I checked” is a sentence, and sentences come from the odds.

THE MODEL WRITES TEXT. THE APP DOES THE THING. THE WINDOW GETS THE RESULT.

The model asks, the app executes, the window receives. The model itself never touches a calculator, a browser, or a file — it writes text that makes the app do the thing, then reads the text that comes back. In the first lesson's x-ray you even saw the standing instruction for this — it is still there to reread: the TOOLS block tells the model to prefer running code over doing arithmetic itself.

The move this buys you: distinguish performed from narrated. A model can write "I ran the numbers" without any tool having run — that sentence is a likely continuation like any other. The app's visible tool step is the receipt, and it is worth knowing its face: a receipt is a step the app renders — a labeled tool run, usually expandable to show what executed and what came back. Code printed inside the reply is not a receipt; it is more sentences, and a stated "output" underneath it is the same fluent guess in costume. When digits matter — an invoice, a forecast, a payroll run, a deadline — make the computation run where you can see it ran: ask for it as code and look for the receipt, or run the arithmetic yourself. No receipt, no computation.

The drowned fact

The second shape you have half-met already. The window is finite, and everything in it competes — you priced that in the first lesson. Here is what the first two lessons did not say: even inside the budget, the pile is not a filing cabinet. Where a fact sits, and what surrounds it, change how reliably the continuation uses it.

Predict again: a contract's liability cap sits on page 2 of a 120-page bundle you pasted whole. Same cap on page 60. Same cap on the last page. Which seat does the reply use most reliably?

TOY — LOADING…

Read the demo's claim carefully, because this is measured terrain and the measurements move: in tested tasks, models often used a fact near an edge of the window better than the same fact buried mid-pile. Often — not always; the effect varies by model, task, and document, and it shrinks as the machines improve. What does not move is the mechanism underneath: retrieval from the window is not lookup. Every token competes for the continuation's attention, and a fact drowning in text that resembles it — ten near-identical clauses, eight drafts of the same table — competes hardest of all. Flip the demo's second control and watch it from the fact's side: the seat doesn't change, but the needle stops standing out from pages that look like it.

Practitioners call the lived version context rot: the longer and messier a window gets, the less reliably any single thing in it gets used. The thread that "got dumb" after an hour did not get tired. Its window filled with dead weight, and your newest sentence became one voice in a very large room.

Three moves fall straight out:

  • Put the load-bearing fact where it can be seen — near an edge, the start or the end of what you paste, not buried mid-pile — and say that it is the load-bearing fact.
  • Prune the lookalikes. The buried fact's worst enemy is similar text it does not need beside it. Delete the haystack you aren't using; don't just relocate the needle.
  • When a long thread has gone dumb, stop pushing it. Start a fresh window and carry a clean summary forward. You know from the first lesson why the machine will not miss the dead weight: nothing runs between sends, and the window you build next is the only memory there is.

The pull of the average

Third shape — the one you will meet most often without recognizing it: the distinctive ask that comes back bland.

You wrote a sharp request — your voice, your angle, a real opinion — and the reply reads like everyone's newsletter. No malfunction. Bland output is the mechanism working: the machine returns the statistically expected continuation of the window you built, and the statistically expected continuation of most professional writing is hedged, balanced, and shaped like the middle of its genre. Your one distinctive sentence was outvoted by everything else in the pile. The shaping you read about in the origin story pushes the same direction — the example answers it learned from were written to be helpful, balanced, and complete, and that polish lives near the middle of the road. Two currents, one drift: toward the average.

The move: move the average on purpose. Name the audience. Force the format. Ban the hedge you don't want — an instruction like end with one recommendation, not a balance of considerations is window text voting against the genre's middle. Paste an example of the register you want — an example in the window is the strongest vote you can cast against habits from training. And when the draft still comes back generic, audit the window before you run it again: what else did you load that votes for generic — the boilerplate brief, the safe template, the average example?

Too many instructions

Now the same mechanism from the instructions side — because the obvious fix for everything above is "add more rules," and that has a failure shape of its own.

You watched this in the first lesson without naming it. The x-ray's instructions block — still there to check — said if the user asks for something the document cannot support, say what is missing, and the memo-shaped summary won the draw anyway. Instructions tilt the odds; they don't command.

So stack constraints and watch what tilting does. Every rule you add narrows what counts as a good continuation. Add enough at once and no continuation scores well on all of them — and nothing forces the machine to stop, error, or ask which rule to break. It writes the best-scoring continuation available and quietly sacrifices the rules that fit worst. There is no error signal for a violated instruction. The reply arrives fluent, confident, and missing constraint number eleven — and nothing in it tells you that, either. In tested tasks the slope repeats: the more simultaneous rules a request carries, the more of them get silently dropped — how steep varies by model, and it moves with every generation.

Two moves:

  • Fewer constraints, ranked. Decide which rules are load-bearing and say so — the figure must come from the attached sheet outranks keep the tone warm, and the window should know it.
  • Split the passes. Draft under the few rules that matter most. Revise under the next few. Check the survivors each pass — a dropped rule you catch in round two cost you a minute, not the deliverable.

Notice this is the instructions-side twin of the balance the second lesson taught you: enough context that it isn't guessing, little enough that nothing drowns — enough instruction that it aims true, few enough that none get dropped.

The world ends on a date

A quick shape, easily named now that you own the two-source ledger. Everything baked in was collected up to a date — the industry calls it the knowledge cutoff — and coverage thins on the approach: training does not sample the world's final months as densely as the years before them, so the built-in world does not end at a wall so much as fade toward one. Treat the published date as a warning line, not a guarantee on either side of it. Anything newer — the price, the version, the ruling, the person's new title — can only arrive through the second door: the window. You even watched the plumbing for this: the x-ray's window carried a supplied date block, because the model does not know today's date until the app pastes it in.

The move follows from what you already do. For anything that could have changed — prices, people, products, law — either watch the search step run (the receipt again) or paste the current fact yourself. If neither happened, the answer was drawn from the fading world behind the warning line. Each model's cutoff is a lookup in its product's documentation — asking the model itself for its cutoff produces, of course, a completion — right only when the real date was pasted into the window or deliberately trained in, and you cannot tell which from your chair. The documentation is the checkable source.

The shape of the hole

You already own the definition: a hallucination is fluent completion with nothing under it in the window — you watched one summarize a memo that wasn't there. What the definition doesn't tell you is what to look for, because a hallucination is not a glitch with a look. It takes the shape of the hole it fills.

And the hole is not always an empty window. A source can be present and still lose the draw — buried mid-pile, outvoted by lookalikes, or simply misread — and the fill looks exactly the same from your chair. Presence in the window is eligibility, not use; you watched that mechanism two shapes ago. Which is why the checks below don't ask how the hole opened.

Ask for sources and the hole is citation-shaped, so the fill is a citation: authors that sound real, a journal that exists, a year in range, page numbers. Ask for precedent and you get a case name shaped like case names. Ask for code and the function is named exactly the way the programming toolkit it claims to come from would name it — if it existed. Ask for the quarter's figure and the number is plausibly round and plausibly sized. The completion matches the pattern of the request, not any source — and by now you can say why: the odds produce what usually follows a request like yours, and what usually follows is an answer-shaped answer.

So verification is not one habit; it is a shape-match:

  • Citations get resolved. Open them. A reference that cannot be opened is a fill until proven otherwise.
  • Quotes get string-searched. Exact words, present in the source, or it is not a quote.
  • Numbers get recomputed — through the tool path, or by hand.
  • Code gets run. A function that reads right and calls a method that does not exist is the same fill in a different costume.

And the discipline that catches the whole class early: give absence its own word. You learned it in the first lesson as UNASSIGNED for owners; the general form is UNKNOWN. An instruction like where the window does not contain it, write UNKNOWN makes not-knowing a completable pattern — an honest continuation that can win the draw. It tilts rather than commands; you watched the honest branch lose once with the instruction present. But over many runs, a window with a word for absence beats a window without one — that is the difference between a report you can audit and one you have to re-derive.

The countermeasure layer

Step back and look at what you are now holding, because it is more organized than it felt on the way in.

Every shape in this lesson is the odds meeting the window you built. The wrong number is computation asked to come out of odds. The drowned fact is a window loaded past what attention can carry. The bland draft is a window outvoted by its own filler. The dropped rule is a window asked to satisfy more than any continuation can. The stale fact is a window that never received the update. The confident invention is a window with a hole where the source should be.

Which is why every countermeasure is a window move:

THE COUNTERMEASURE LAYER — MATCH THE SHAPE TO THE MOVESCHEMATIC — THE SHAPE, NOT THE SIZES
  • WRONG NUMBERno computation you can audit — digits from the odds
    COMPUTE, DON'T COMPLETEroute it through a tool; demand the receipt
  • DROWNED FACTburied mid-pile among lookalikes
    CURATE THE WINDOWload less, near an edge; fresh window, clean summary
  • BLAND DRAFTthe genre's middle, everyone's draft
    MOVE THE AVERAGEname the audience, force the format, ban the hedge
  • DROPPED RULEconstraint eleven silently gone
    SPLIT THE PASSESfewer rules, ranked; revise in rounds
  • STALE FACTcoverage that thinned out toward the cutoff
    FEED THE WINDOWwatch the search receipt, or paste the update
  • CONFIDENT INVENTIONa fill in the hole's own shape
    GIVE ABSENCE A WORDUNKNOWN in the window; then resolve, recompute, rerun
EVERY SHAPE TRACES TO THE ODDS MEETING THE WINDOW. EVERY COUNTERMEASURE IS A WINDOW MOVE.

Compute instead of complete. Curate instead of dump. Move the average on purpose. Split instead of stretch — fewer rules per pass, one job per window, clean summaries carried forward, because a window that does everything fills with everything. Give absence a word. And judge from a second window — the first lesson's discipline, still the cheapest audit there is. None of this needs a better model, a different vendor, or a setting you haven't met. It is all pile management: the craft the second lesson named minimum viable context, now with the failure shapes that explain why it works.

One more time, because it is the honest limit of everything above: these are the shapes, not the map. Where the edge runs through your work — which of your tasks it drafts brilliantly, which it quietly fumbles — cannot be deduced from here, by you or by anyone. The labs are how you map it: real work, on your own material, where you can judge output cold and now know exactly what to watch for.

Check your understanding: a colleague pastes a 90-page bundle into a chat, adds fourteen rules, asks for the exact liability figure, and gets a fluent, confident, wrong answer. Name the three shapes that just stacked — and the window move that answers each one.

Check yourself first — then open one good answer

The stack, named: the 90-page paste is the drowned fact — the liability figure buried mid-pile among clauses that resemble it. Fourteen rules at once is too many instructions — some quietly sacrificed, no error signal. And an exact figure asked from the odds is the wrong number waiting to happen — answer-shaped digits, typographically identical to a computed result. (If the figure was never in the bundle at all, the same request breeds a fourth shape: a number-shaped confident invention.)

The moves, matched: curate the window — load the section that matters near an edge and prune the lookalike drafts. Rank the rules — the figure must come from the attached document outranks the rest, which can wait for a second pass. Compute, don't complete — route the extraction through the tool path and look for the receipt, or pull the figure by hand.

Your answer should have hit the three shapes — drowned, dropped, wrong number — each answered by its window move, and no blame on the machine.

Use this understanding as your checklist. Every lab from here on runs at this edge. When an answer disappoints, name the shape before you blame the machine: computed or completed? drowned or seen? averaged or steered? dropped or ranked? stale or fed? filled or sourced? The shape names the fix, and the fix is always a window move. Next stop — the labs: the map-making hours, arranged.

← All labs