Answer
How do you use AI for research without being misled?
Treat the output as a list of places to look rather than as findings, and open every source before relying on it.
Treat it as a route to sources rather than a source. Every claim is a pointer to something you should open, and the research is what you find there rather than what the summary said.
The strength here is coverage and speed: locating relevant material, identifying what has been written on a question, surfacing sources you would not have found. The weakness is that the account of what those sources say is a summary produced without you seeing them, subject to the ordinary asymmetry that omissions do not announce themselves. Treating the account as the research is where people get into trouble.
So the discipline is to convert the output into a reading list. Each claim should carry the source it came from, each source should be opened, and the claim should be checked against what the source actually says. This sounds laborious and is quick in practice, because opening a page and finding a sentence takes seconds, and because most claims turn out to be fine. The ones that do not are the reason to do it.
Three failures recur and are worth watching for specifically. A source that does not contain the claim attributed to it, which is common enough to be the default suspicion. A source that contains it with a qualification that did not survive — a date, a population, a condition. And a claim attributed to a secondary source that was itself repeating something, where the trail leads to another summary rather than to a measurement.
That last one is the most consequential for anything you will publish or act on. A figure repeated widely acquires the appearance of established fact, and the repetition is frequently one original claim carried by many outlets. Following it back to whoever actually measured is the only way to know what it means, and the exercise regularly ends at a source that measured something narrower than the way the figure is used.
Absence of evidence needs its own handling because it is the least reliable output. If a search finds nothing, that may mean nothing exists, or that it was not looked for effectively, or that it exists under different terminology. A system reporting that no source addresses a question is making a claim about its search rather than about the world, and the two are distinguishable only by asking what was searched.
One thing that genuinely helps: ask for the search rather than the answer. Which terms, which sources, what was found and what was not. That output is checkable, it tells you whether the coverage was adequate, and it turns the exercise into something you are directing rather than receiving. It also surfaces the case where the search was narrow, which no confident summary reveals.
It is a very fast way of finding out where to look, and a poor substitute for having looked.
Siddharth Sharma, Context Theory
Related questions
Are deep research modes more reliable?
They cover more ground and produce a longer account, which improves the sourcing and does not change the verification requirement. A longer document with more citations contains more claims to check rather than fewer, and the length can make the checking feel unnecessary, which is the specific risk it introduces.
What about subjects where you cannot evaluate the sources?
That is the case to be most careful about, because you are relying on the summary of material you could not have assessed anyway. What remains available is checking that the sources exist, that they are the kind of source they claim to be, and that they say what is attributed to them, which is worth doing and is less than understanding the field.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| AI-cited sources that also rank in the Google organic top 10 | 10% | Category-wide |
| Visibility lift in AI-generated answers from GEO methods | up to 40% | Category-wide |
2026 generative engine citation study · fewer than · verified
Aggarwal et al., "GEO: Generative Engine Optimization", Princeton / Georgia Tech / IIT Delhi / Allen Institute for AI — KDD 2024 · GEO-bench · 10,000 queries across 8 domains · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | The reliable contribution is locating material rather than characterising it, because the account of what a source says is a summary produced without the reader seeing the source, subject to unannounced omission. | Opening the sources cited in a research output and comparing each against the claim attributed to it. |
| Response | Three failures recur: a source lacking the attributed claim, a claim stripped of a qualifying date, population or condition, and a claim attributed to a secondary source that was itself repeating something. | Classifying discrepancies found when checking a research output against its cited sources. |
| Software | A widely repeated figure frequently originates in a single claim carried by many outlets, and tracing it back regularly ends at a source that measured something narrower than the way the figure is used. | Following a commonly cited statistic through its citing sources to the original measurement. |
| Constraint | A report that no source addresses a question is a claim about the search rather than about the world, and the two are distinguishable only by asking what terms and sources were used. | Requesting the search terms and sources consulted alongside any not-found conclusion. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one