Practice Lab

Written to Be Retrieved

EssayAugust 20, 2026 · David J.S. Madgett · 17 min read

Somewhere in Minnesota tonight, a person with a legal problem and no money is going to type it into an assistant instead of calling a lawyer. That has been true for a couple of years. What is newly true is what happens next: the assistant does not hand back ten blue links for the person to evaluate. It composes an answer, and to compose it, it goes and takes passages off pages like mine.

That changes who I am writing for, and I want to be precise about the change rather than gesture at it, because the gesture — “optimize for AI!” — has already produced an entire consulting industry selling the same thin content it sold last decade under a new three-letter acronym.

Here is the thesis, stated so it can be checked: the reader that matters most for a law firm’s published writing is now a retrieval system, and a retrieval system rewards depth, verbatim accuracy, pinned authority, stable addresses, and explicit statements of what an authority does not cover. Which is to say it rewards the things that make legal writing good, and punishes the things that made SEO content bad. Writing for machines that read on behalf of humans pushes you back toward genuine quality. That is a happy accident and it is the most interesting thing about this shift.

Search asks which page. Grounding asks what can be said.

You do not have to take this from me. The clearest statement of the change I have read was published by Bing’s own engineering team in May 2026, and it is worth reading in full if you publish anything.

Their framing: traditional search asks “which pages should a user visit?” Grounding an AI answer asks “what information can an AI system responsibly use to construct a response?” Those sound like the same question. They are not, and the post lays out the divergence in a table. In traditional search, the “unit of value” is the document (page). In grounding, it is “groundable information (discrete, supportable facts with clear provenance).” In traditional search, “imperfect ranking is tolerable; recovery is easy,” because a human scans the results, skips the bad ones, and course-corrects. In grounding, “errors can compound across reasoning steps.”

Then the line that should stop any lawyer reading it: among the valid outcomes for a grounding system, Microsoft lists “answer when supported; abstain when evidence is insufficient.” And on freshness — in search, “stale content degrades ranking”; in grounding, “a stale fact produces a misleading response.”

Read that as a lawyer and it is familiar. A system that must decide whether the evidence supports the assertion, that must flag conflicting sources rather than silently pick one, that must know when to say nothing — that is not a search engine’s value system. That is a research memo’s value system. Bing’s own summary of the shift: “search indexing was built to help humans decide what to read. Grounding indexing is being built to help AI systems decide what to say.”

If the machine reading your page is evaluating whether your sentence is safe to repeat, then everything about what makes a page valuable inverts.

The new reader is a hostile cite-checker

Work through what that reader actually wants, and it is a list every litigator already knows.

Provenance. Microsoft’s version: grounding requires “high-quality source identification and attribution so users (and downstream systems) can verify what was used and follow the evidence when needed.” Ours: a pin cite. A page that says “under Minnesota law, you generally have six years” gives a retrieval system nothing it can stand behind. A page that says which section, which subdivision, and quotes the operative clause gives it something it can commit to and attribute.

Factual fidelity through chunking. This is the mechanic worth understanding, and Microsoft is blunt about it: the processes of “breaking content into retrievable chunks and transforming it for fast lookup can distort page substance in ways that never appear in any ranking signal.” Your page is not read; it is cut into pieces and the pieces are ranked. So a paragraph that depends on three paragraphs of setup to be accurate is a paragraph that will eventually be quoted wrong. The corollary is a drafting rule: each paragraph should survive being lifted out of the page. Microsoft’s own guidance says the same thing from the other direction — content is eligible for selection when it has “self-contained phrasing: sentences that make sense even when pulled out of context.”

Lawyers have a name for a sentence that changes meaning when separated from what precedes it. We call it a quote taken out of context, and we spend our careers objecting to it. Now the retrieval layer does it to every page on the internet, mechanically, at scale, and the only defense is writing paragraphs that are true standing alone.

Conflict. A grounding index “must detect and represent conflict; silent arbitration risks confident wrong answers.” This is the part I find most encouraging, because conflict is where legal writing actually lives. The published Court of Appeals decision that cuts against the trend, the 2026 amendment that has not made it into the codified text yet, the split between districts — a thin content farm flattens all of that into a confident sentence. A page that says “the Court of Appeals has said X; the supreme court has not addressed it” is more useful to a system trying to decide what it can responsibly assert, not less.

Thin content was always bad. Now it is invisible.

The old game rewarded a 600-word page titled “What Is a Personal Injury Claim in Minnesota?” that never cited a statute, because ranking was a competition between documents and a document could win on keywords, links, and freshness signals without containing anything.

Under grounding, that page has nothing groundable in it. There is no discrete supportable fact with clear provenance to extract. It can rank and still never be cited, because the thing being selected is not the page.

The alternative is expensive and I can only describe what it costs by describing what this firm actually built. As of today the news library on this site is 321 published articles running to roughly a million words. Three hundred of those 321 contain at least one citation to a Minnesota statute; the corpus contains 3,093 separate references to “Minn. Stat.” The statutory quotations are pulled from the Revisor’s raw HTML rather than through a summarizing fetch layer, for reasons I have written about separately — a summarizer paraphrases, and a paraphrase inside quotation marks is a fabrication. Case quotations are read from the opinion, with pin cites resolved against star pagination. Every article closes with a sources block naming each authority and the exact proposition it supports, down to the page.

That is not a content strategy. It is the ordinary standard of care for legal writing, applied to publishing. The reason it is worth mentioning at all is that for ten years the market rewarded not doing it, and the market has changed its mind.

Say what the statute does not say

This is the specific move I would push hardest on, because it is the one that thin content structurally cannot make.

Search across this site’s articles and you find the word pattern everywhere: 22 places that say a statute “does not define” a term, 80 that say some authority “does not say” something, three that describe a statute as silent, four that report finding no Minnesota case on a point. One article works through the honesty-testing statute and stops to note that the prohibition reaches “any test purporting to test the honesty” of an employee, that the statute does not define “test,” and that the phrase therefore names a purpose rather than a technology. Another reports that the Court of Appeals “found no Minnesota case allowing recovery in the absence of direct physical injury to the spouse.”

To a human skimmer, admitting a gap reads as weakness, which is exactly why marketing copy never does it. To a system deciding whether the evidence is sufficient to answer, a clean statement of the boundary of an authority is among the most valuable sentences on the page. It is the difference between a source that lets the assistant answer confidently within a limit and a source that invites it to overstate.

It is also, not incidentally, how you talk to a client who deserves to know that the answer is unsettled.

A URL is a citation format

If your page is going to be cited, its address has to behave like a citation.

Articles here live at dateless slugs — /news/<slug>/ — with the publication date carried in frontmatter rather than in the URL, so that revising a guide in 2027 does not require moving it. When something has had to move anyway, a 301 goes into the redirects file; there are 41 rules in it now, most of them pointing old addresses at current ones.

The subtler piece is lastmod. The sitemap integration will happily stamp every URL with the build time, and the comment I left in the config when I fixed it says why that is a problem: it “tells a crawler that 212 articles changed every time one of them is edited — and it learns to stop believing the file.” So the sitemap’s lastmod is read out of each article’s own frontmatter — the updated date when a piece has been revised, otherwise its publication date — and the same date drives dateModified in the article’s schema.org markup. Topic hub pages carry CollectionPage and ItemList schema listing their members, so a crawler sees nine subject clusters rather than 321 unrelated pages, and every article’s publisher points at a single firm entity node declared once rather than restated 321 times.

None of that is clever. It is bookkeeping. But freshness is one of the places where grounding and search genuinely differ in consequence — stale ranking is an annoyance; a stale fact in a composed answer is a wrong answer with your name on it as the source — and a lastmod field a crawler has learned to ignore is worse than none.

The plumbing, stated precisely

Here is where I want to be careful, because this is the part of the subject where everyone overclaims.

IndexNow is a real protocol with published mechanics. You host a text file at your root whose name and contents are both your key; you POST JSON with host, key, and a urlList to a participating endpoint; you may “submit up to 10,000 URLs per post.” A 200 means accepted, a 202 means received with key validation pending, a 403 means the key did not check out. The protocol’s own documentation is careful about what a success means: “The HTTP 200 response code only indicates that the search engine has received your URL.” Participating endpoints listed at indexnow.org today are the IndexNow global endpoint, Amazon, Bing, Naver, Seznam.cz, Yandex, and Yep — and submissions to any one of them are shared with the others. Google is not among them.

Deploying this site ends with a script that submits the new URLs. It checks that the key file is actually reachable before submitting anything, because IndexNow rejects the batch if it cannot fetch the key and the failure is otherwise silent.

What it buys you is the part to state exactly. Microsoft says that IndexNow “helps ensure that AI systems reference the most current version of a page when generating answers,” and separately that “Powered by Bing’s search index, experiences like Microsoft Copilot, Microsoft Start, and others handle billions of queries each month.” Microsoft’s new AI Performance dashboard reports citations across “Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations.”

Note what Microsoft does not say: which partners. You will read confident assertions that Bing’s index is what backs this or that third-party assistant. I cannot verify that from a primary source and so I am not going to write it. The verified claim is narrower and still worth acting on: pinging IndexNow gets a new page into Bing’s index in minutes rather than weeks, and Bing’s index is, on Microsoft’s own statement, what Copilot and Microsoft Start run on.

The other assistants crawl for themselves, which makes robots.txt an editorial decision rather than a technical one. OpenAI’s documentation is explicit: OAI-SearchBot “is used to surface websites in search results in ChatGPT’s search features,” and “sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.” Anthropic documents three separate agents — ClaudeBot for training data, Claude-User for user-initiated fetches, and Claude-SearchBot for search indexing — and says that disabling Claude-User “prevents our system from retrieving your content in response to a user query.”

This site allows all of them, and says so by name in robots.txt. Worth noticing, though: that named list is already stale. Claude-SearchBot is not on it, because it was documented after the file was written. The only reason that does not matter is that the wildcard rule allows everything, which is what is actually doing the work. A list of named crawlers in a robots file is a statement of intent; the default is the policy.

I publish an llms.txt and I do not yet believe in it

Honesty requires a section on this one.

The llms.txt proposal — Jeremy Howard’s, first published September 2024 and now in a second version modified August 10, 2026 — asks sites to serve a markdown file that gives an agent a curated map: an H1 with the site name, a blockquote summary, then H2 sections containing lists of links with short descriptions. The reasoning is sound. Web pages “wrap information in navigation, ads, and JavaScript,” conversion back to clean text is “difficult and imprecise,” and “every wasted token costs time and money.” The design intent is that the file itself “stays small enough to fit in context,” with detail behind the links.

This site generates one at /llms.txt from the content collections at build time, plus a /llms-full.txt carrying the source markdown of every published page.

Two things I will not pretend about.

First, the evidence for adoption is all on the publishing side. The spec page documents thousands of sites publishing the file, documentation platforms generating it automatically, and Chrome’s Lighthouse auditing for one. It documents no assistant vendor committing to read it. Google’s John Mueller, asked about it in June 2026, called it “purely speculative for now,” and added the observation that lands: “the file has existed for years, yet none of the AI systems use it.”

Second, my own implementation violates the spec’s central premise. /llms.txt here is 158 KB and /llms-full.txt is 6.9 MB. Neither of those is “small enough to fit in context” in the sense the proposal means. And the v2 spec’s stronger recommendation — serving a clean .md twin of each page at the same URL, with rel="alternate" and rel="describedby" link relations pointing to it — is something this site does not do at all.

So the honest posture is: it costs nothing to generate, it is generated correctly enough, I have told you it is unproven, and if I am wrong about it I have wasted forty lines of build code. That is the whole case for it. Anybody selling you an llms.txt engagement is selling you something else.

Nobody can measure this yet

The most important thing I can tell another lawyer about writing for retrieval is that you cannot currently tell whether it is working, and you should distrust anyone who tells you they can.

Assistant citations do not show up in your analytics as referrals in any consistent way. The one real instrument I know of is Bing Webmaster Tools’ AI Performance view, released in public preview in February 2026, which reports total citations, average cited pages per day, page-level citation counts, and the “grounding queries” — “the key phrases the AI used when retrieving content that was referenced.” Microsoft attaches its own caveats to every one of those metrics: the citation count “does not indicate placement or presentation within a specific answer”; cited-page averages do “not indicate ranking, authority, or the role of any page within an individual answer”; the grounding-query data “represents a sample.”

And that instrument covers Microsoft’s surfaces plus unnamed partners. There is no equivalent for ChatGPT, Claude, or Gemini that I can point you at. Bing’s own engineers say it plainly: “What makes this hard is not the technology gap — it is the measurement gap.”

So I am not going to show you a chart. I am going to tell you what I built, why the reasoning holds, and what would falsify it.

What would prove me wrong

The thesis fails if what gets selected into AI answers turns out to be short, structured, thin content rather than depth with pinned provenance — if a page of Q&A blocks with no authority outperforms a three-thousand-word piece that quotes the statute, because the extraction layer rewards format over substance and never gets far enough into the page to notice the substance.

That is not a strawman, and Microsoft’s own optimization guidance contains the tension. It says to “avoid long walls of text” because they “blur ideas together and make it harder for AI to separate content into usable chunks,” and it recommends lists, tables, and Q&A pairs, which “assistants can often lift word for word.” It even advises being “cautious with em dashes,” on the theory that overuse “can confuse sentence structure for machines.”

I am not giving up em dashes. But the substantive half of that guidance is a real correction to the thesis and I would rather absorb it than argue with it: the reward is for structure, not for length. Depth without structure loses to shallowness with structure. A long undifferentiated essay is as hard to chunk as a short empty one. Headings that name the actual question, one idea per paragraph, the quotation set off where it can be lifted cleanly, the sources named at the foot — that is the format, and it is not a concession, because it is also how you write something a busy lawyer can use.

If, three years from now, the assistants are reliably citing content with no authority behind it and ignoring the primary-source work, I will have been wrong about what the retrieval layer selects for, and the correct move will have been schema tricks and volume. I do not think that is where this goes, because the entire direction of the engineering — abstain when unsupported, detect conflict, attribute so the user can verify — is toward evidence rather than away from it.

The standard is whether an assistant that checks would keep citing you

There is a symmetry here that I think decides the question.

I have written elsewhere on this site about how this firm’s own assistant does legal research: it pulls statutes as raw HTML from the Revisor rather than through a summarizer, reads opinions rather than snippets, and runs a separate verification gate before anything is used. An assistant built that way does not need a law firm’s blog post about a statute. It has the statute.

Which means the only content worth publishing is content that survives being checked against the primary source by a reader who has the primary source open. A page that paraphrases the statute loosely is worse than useless to that reader — it is a liability the moment it is compared. A page that quotes the operative clause correctly, cites the subdivision, notes that a 2026 amendment changed the effective date, and says which question the statute does not answer is a page that reader keeps.

That is the standard. Not “will this rank.” Would an assistant that checks the statute keep citing you.

And this is where the arithmetic of the whole section comes back around. The person who cannot afford a consultation is going to ask an assistant anyway. What that assistant repeats back to them is going to be assembled out of whatever is on the open web about Minnesota law. Right now a great deal of what is out there is lead-generation copy that was never checked against anything, and it is being laundered through a confident interface into what sounds like advice.

There is no gatekeeper who is going to fix that. The only mechanism available is that lawyers who know the law put accurate, cited, honestly-bounded versions of it where the retrieval layer can find them — which costs the firm nothing but the work, and is worth more to a person with no money than any amount of free-consultation marketing.

That is why everything in this section is published free, and it is the same reason to write the guides carefully. The audience was always people who could not afford to ask. Now there is a machine standing between you and them, and it turns out the machine has excellent taste.


Sources

  • Bing Search Blog, Evolving role of the index: From ranking pages to supporting answers (May 6, 2026) — traditional search vs. grounding comparison table; “groundable information (discrete, supportable facts with clear provenance)”; abstention as a valid outcome; freshness and conflict-detection differences; chunking distortion; “the measurement gap”
  • Bing Webmaster Blog, Introducing AI Performance in Bing Webmaster Tools Public Preview (Feb. 10, 2026) — citations across “Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations”; total citations, average cited pages, grounding queries, page-level citation activity, and Microsoft’s caveats on each; IndexNow and answer freshness
  • Microsoft Advertising, Optimizing Your Content for Inclusion in AI Search Answers (Oct. 8, 2025) — “Powered by Bing’s search index, experiences like Microsoft Copilot, Microsoft Start, and others”; parsing into modular pieces; self-contained phrasing; guidance against walls of text, hidden content, and PDF-only information; the em-dash caution
  • IndexNow protocol documentation — key-file ownership verification, POST JSON format, 10,000-URL batch cap, 200/202/400/403/422/429 response semantics, and the requirement that participating engines share submissions
  • IndexNow FAQ — the list of participating endpoints (IndexNow global, Amazon, Bing, Naver, Seznam.cz, Yandex, Yep); Google is not among them
  • Jeremy Howard, The /llms.txt file, v2 (published Sept. 3, 2024; modified Aug. 10, 2026) — format specification, the “small enough to fit in context” design goal, the v2 markdown-twin and link-relation recommendations, and the publisher-side adoption evidence
  • Search Engine Journal, Google Says LLMs.txt Is “Purely Speculative” For Now (June 2, 2026) — John Mueller: “purely speculative for now (the file has existed for years, yet none of the AI systems use it)”
  • OpenAI, Bots — OAI-SearchBot surfaces sites in ChatGPT search features; “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers”; GPTBot and ChatGPT-User distinguished
  • Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated Apr. 7, 2026) — ClaudeBot, Claude-User, and Claude-SearchBot, and the consequence of disabling each
  • Site figures (321 published news articles, ~1,000,000 words, 3,093 “Minn. Stat.” references across 300 articles, 22 “does not define” passages, nine topic hubs, 41 redirect rules, 158 KB /llms.txt, 6.9 MB /llms-full.txt) were counted from this site’s own repository and live URLs on August 20, 2026.

Commentary on legal publishing and practice technology; the opinions and predictions are the author’s, offered with failure conditions so they can be tested rather than relied upon. Not legal advice, not ethics advice, and not a recommendation of any vendor or optimization service. Nothing here reports measured results — as stated above, the measurement tooling for assistant citations barely exists. No client information appears in this article. Questions about anything here: Send us a message or 612-470-6529.

words
3,975
sections
11
sources
9
distinctive_terms
microsoft · bing · grounding · indexnow · content
Pass it onLinkedInX

Get new articles as they land

One email when something new is published here. No course, no upsell — the Practice Lab stays free either way.

Used only to send Practice Lab posts. Unsubscribe from any email. Subscribing does not create an attorney–client relationship.

The only thing we ask

If something here saves you time, spend some of it on people who could not otherwise afford you.

Everything in the Practice Lab is free. No signup, no subscription, no donations — just take a case you would otherwise have to turn down on economics. More from the Practice Lab →

Keep Reading

12% vocabulary overlap

Build a Firm Wiki Your AI Can Query

The reason AI gives you generic legal work is that it has no idea where it is. A structured Obsidian knowledge base — plain markdown, real schema, provenance on every fact — turns a general assistant into one that already knows your practice.

Workflow · 15 min read

11% vocabulary overlap

Retrieval Is the Whole Game

Why dumping a matter file into a context window fails, why embedding search underperforms specifically on legal text, and how to build retrieval that recovers the right paragraph with a citation attached.

Tool · 18 min read

10% vocabulary overlap

The Commons Your Practice Runs On

A solo firm can now verify law like a research department, and the reason is a handful of nonprofits and public offices that almost nobody in this profession pays — so here is the honest structural survey of that infrastructure, the four specific places it is already breaking, and the argument for a donation line item that is not charity but a maintenance bill.

Essay · 17 min read

← All Practice Lab articles