AI tools for investigations are research agents that start from an identifier, such as a name, an email address, a username, a phone number or a company, and search public sources to build a sourced picture of who or what is behind it: accounts, records, connections and adverse media. They take on much of the collection and cross-referencing an analyst would otherwise do by hand. Before relying on one, check five things: that every finding carries a source you can open, that it says "unknown" instead of guessing, that it tells apart people who share a name, that its collection stays passive, and that it gets cases right when you already know the answer.
By investigations, this piece means people and company investigations: trust and safety, fraud, due diligence, hiring and open-source intelligence work.
What these tools do well
Much of an investigation's time goes on collection: searching one identifier across dozens of sources, following it to the next identifier, reading records in other languages, and noting where each fact came from. That work is repetitive, and it grows with every identifier you follow.
This is the part an agent is good at. It can run many searches at once, pivot from an email address to a username to an old forum account without losing the thread, read a scanned filing, and hand back a structured result with its sources attached. Our guides to investigating an email address and spotting fake candidates show what that collection involves when it is done by hand.
The judgement stays with the investigator: whether a link is strong enough to act on, what a finding means for the case, and what happens to the person or company at the end of it. The agent collects and organises, and the investigator decides.
Five checks before you rely on one
1. Every finding has a source you can open
Citations only help if you open them, because a sourced finding can still be wrong. Check that each page says what the tool says it does.
A 2025 study involving 22 public service media organisations shows why. It found serious sourcing problems, meaning missing, misleading or incorrect attributions, in 31 percent of the answers that four widely used AI assistants gave to news questions. Those were general assistants rather than investigation tools, but the failure is the same one to test for.
Test it directly. Take twenty findings from a few results, open every source, and count how many support the claim as written. Watch for the near misses: a real page that mentions a different person with the same name, or an article that says something weaker than the summary above it.
Sources need dates as well, because pages change and come down, and a finding is only as current as the day its source was read. The Berkeley Protocol on Digital Open Source Investigations, published by the UN human rights office and UC Berkeley's Human Rights Center, asks human investigators to record the search paths that led to each piece of material, when and by whom it was collected, and their method in enough detail that someone else could repeat it. A tool that does the searching for you should keep that record for you.
2. It says "unknown" instead of guessing
For an investigator, a blank field is cheap and a wrong one is expensive. A blank tells you to keep looking. A confident wrong answer can end up in a report.
Language models lean the other way by default. A 2025 paper on why language models hallucinate argues that standard training and evaluation reward a confident guess over an admission of uncertainty. The broadcasters' study above found major accuracy issues, including hallucinated details and outdated information, in 20 percent of the answers it reviewed.
So test for restraint as well as recall. Include a few subjects with thin public footprints and see whether the tool returns gaps or fills them. The result you want is a list of what was established, what was not, and why.
3. It gets the right person among people who share a name
In the 2010 US Census, 2,442,977 people had the surname Smith, and each of the ten most common surnames was shared by more than a million people. A wrong identity match is one of the most expensive mistakes in people research, because everything found after it belongs to someone else.
A good tool anchors to what you already know, such as an employer, a city, an email address or a date of birth, and explains why each new account or record belongs to the same person. Test it with a subject whose name is common, and with one who has a namesake in the same field or city. Read the reasons it gives for each link, not only the links.
4. Collection stays passive
Open-source collection should not touch the subject. The Berkeley Protocol draws its definition of open source information along that line: such information "does not involve interacting with or soliciting information from individual Internet users". For a tool, that rules out logging into the subject's accounts, triggering password resets, messaging people, or anything else that would tell someone they are being looked at or that relies on access nobody authorised. Ask the vendor directly what the tool does to collect, and look for anything in the output that could only have come from an active step.
Ask where your queries go as well. The names and identifiers you search are sensitive in their own right, and you should know how long they are kept and who can see them.
5. It holds up on cases you already know
Your own closed cases make the best test, because you know what the right answer was and how hard it was to find. The next section sets out how to build one.
Run the same case twice as well. Systems built on language models can give different answers to the same question even with settings meant to make them repeatable, and a 2024 study of "deterministic" model settings measured variations in accuracy of up to 15 percent across runs on standard benchmarks. Some variation is normal, and it should sit at the edges of a result, such as an extra old account found one time and not the other. The core identity should not change between runs.
Build your own test set
Twenty to thirty closed cases should surface most of a tool's failure modes. Pick them so that each one tests something specific:
| Include | What it tests |
|---|---|
| Subjects with rich public footprints | Recall, and whether sources are cited properly |
| Subjects with thin footprints | Whether the tool reports gaps instead of filling them |
| Common names, and pairs of people who share a name | Identity resolution |
| An identifier that changed hands, such as a recycled phone number or a reassigned work email address | Whether the tool checks dates before linking |
| Names in other scripts or languages | Transliteration and non-English sources |
| A company with a similarly named counterpart | Entity resolution for organisations |
Score every field as correct, wrong or missing, and weigh a wrong answer as worse than a missing one. Separately, count the citations that do not support the claim they are attached to.
The first half of that is how RECON scores. RECON is the people research benchmark we built, externally validated by HUD, and it covers 140 real people and 514 verified fields, with roughly a quarter of the people paired with a same-named decoy. Its headline metric is weighted accuracy: correct answers minus wrong answers, divided by every field. A wrong answer costs a point, and a blank costs nothing.
The most accurate system on the benchmark is our highest tier. In the published results it answered 348 of the 514 fields correctly, answered 69 wrongly and left 97 blank, which is roughly one wrong answer in every six it gave. Every system tested got some answers wrong, so a person should read the result before anyone acts on it. The full results and methodology are published.
What Sixtyfour's agent returns
Sixtyfour builds AI research agents for people and company investigations. The agent starts from the identifiers you have, such as a name, email address, phone number or company, researches across public records, breach records, social profiles, code hosting, forums, archived pages, filings and news, and returns findings with a source attached to each one. It says which accounts and records it believes belong to the same person and why, and reports what it could not establish as open rather than filling it in, so an investigator can check every finding before acting on it.
The checks, in brief
- Open the sources. Count how many support the claim as written, and check they are dated.
- Give it thin footprints and see whether it admits gaps.
- Give it common names and pairs of people who share a name, and read its reasons for each link.
- Ask what it does to collect, and where your queries go.
- Run your own closed cases, twice, and score wrong answers as worse than blank ones.
Run these checks on your own cases and you will know where a tool's results can be trusted and where a person still needs to look.