AODM test results

What AODM tagging changed when AI agents answered questions from tagged and untagged versions of the same pages.

The short version

When an AI works from search results instead of reading a whole page, AODM keeps each fact attached to its source, its date, its unit and its status. Across the tests below, AI agents that searched AODM-tagged pages answered more questions correctly than agents that searched the same pages without tags.

No model invented a value in any run. Without tags, agents mostly said an answer was not in front of them, or repeated something a page claimed without knowing how far to trust it.

Proven benefits

What the AI had to doUntaggedAODM tagged
Say who tested a spec, when and under what conditions, from search results (Apple iPhone 16 specs)64–73%100%
The same, given only the passage holding the claim9%100%
Say which measurements a figure was calculated from17%100%
Point to the exact passage behind an AI-generated claim0%100%
Say when a manufacturer made a claim and when it was retrieved50%100%
Keep benchmark scores tied to the test that produced them once a review is split into passages (GSMArena iPhone 17 Pro review)75%100%
Answer 25 questions covering every AODM feature, from search results (battery lab report)79%88%
Answer questions about an SEC filing from search passages (Apple 10-Q; original XBRL against the same filing in AODM)23%100%
Answer questions through a standard search pipeline over 5 SEC filings from 5 companies (the filings’ HTML against the same filings in AODM)72%88%
Keep each number’s company, quarter, segment, unit and scale attached after 7 SEC filings are split into passages (original XBRL against AODM)0%77%
Pull every segment’s revenue from 7 SEC filings into one table with plain code and no AI (the filings’ HTML tables against AODM)71%100%
Decide 1,000 grant applications the same way on every run, under a policy with dated routes, exceptions and nested exclusions99.8%100%
Say whether a product supports a feature when the comparison table shows it only with icons67%100%
Surface every distinct source when 1 press release is copied across 8 news sites (top 5 search results)65%100%
Keep claims found only on an AI-written page out of answers0%100%
Pick an independent lab measurement over a manufacturer’s rating67%100%
Tell whether a measured value really contradicts a rating, given each figure’s margin of error83%100%
Spot that a recently republished article repeats a 3-year-old claim89%100%
Say “not documented” instead of guessing when no page answers the question17%67%
Trace a 4-step supply chain across 10 products to find which are hit by a factory closure80%99%
Find the report section behind a paraphrased summary claim in a 60-section report61%100%
Compare and total 20 products listed in mixed units98%100%
Give the price that applies on the date asked about, when cached old prices, announced future prices and an expired offer are all in the search index80%100%
Find every indicator a corrected figure affects when the publisher does not explain its formulas0%94%
Remove duplicate copies from a training corpus without losing any unique fact31%75%

Each percentage is the share of correct answers across 3 runs per version. The row on distinct sources measures what the search pipeline returned, and the training row measures the share of duplicate copies removed. For the filings, the untagged column is the original XBRL unless the row says otherwise; for training data, it is the standard exact-match method.

Financial filings show the gap most clearly. XBRL keeps the period and segment of every number in a separate block at the top of the file and links to it by an ID. Once the file is split into search passages, the number usually arrives without them, so the AI cannot tell which quarter or product it belongs to. AODM writes the period, unit and segment onto each fact, so they travel with the number.

Duplicate detection also helps model training. Small language models trained on a corpus deduplicated with AODM fact hashes kept every unique fact, memorised copied text least and predicted unseen text best. A common near-duplicate method lost 35% of the unique facts.

How we tested

  • We used real pages (Yellowstone entrance fees, Apple’s iPhone 16 specs, the GSMArena iPhone 17 Pro review), 10-Q filings from Apple, Microsoft, Alphabet, Amazon, Nvidia, Coca-Cola, Johnson & Johnson, Walmart, Procter & Gamble and Visa taken from SEC EDGAR, and 103 purpose-built pages about fictional companies, so no model could know the answers in advance.
  • Each page had 2 versions with identical visible text: 1 tagged with AODM and 1 with every tag removed. Filings were compared in their original formats against the same data converted with the AODM XBRL converter.
  • In most tests, AI agents had only a search tool over a collection that included distractors: older versions, other quarters, other companies, copied articles and AI-written pages. They wrote their own queries. Search combined keyword and semantic ranking, the way production systems do.
  • We ran Claude Opus, Claude Sonnet and Claude Haiku, 3 independent runs per version, and fixed each test’s questions and grading before running it.
  • Answers were graded for the right value, the right source and date, saying “not documented” when nothing answers the question, and invented values.
  • The training test used small models trained on a Mac with the same compute for every method. It shows direction, not production-scale results.
  • We built the test pages, the tags, the questions and the grading ourselves.