How to evaluate AI tools for investigations

AI research agents can take on much of an investigation's legwork. Five checks tell you whether their results will hold up when someone asks how you know.

AI InvestigationsOSINTAI AgentsDue DiligenceIdentity ResolutionRECON Benchmark
How to evaluate AI tools for investigations

AI tools for investigations are research agents that start from an identifier, such as a name, an email address, a username, a phone number or a company, and search public sources to build a sourced picture of who or what is behind it: accounts, records, connections and adverse media. They take on much of the collection and cross-referencing an analyst would otherwise do by hand. Before relying on one, check five things: that every finding carries a source you can open, that it says "unknown" instead of guessing, that it tells apart people who share a name, that its collection stays passive, and that it gets cases right when you already know the answer.

By investigations, this piece means people and company investigations: trust and safety, fraud, due diligence, hiring and open-source intelligence work.

What these tools do well

Much of an investigation's time goes on collection: searching one identifier across dozens of sources, following it to the next identifier, reading records in other languages, and noting where each fact came from. That work is repetitive, and it grows with every identifier you follow.

This is the part an agent is good at. It can run many searches at once, pivot from an email address to a username to an old forum account without losing the thread, read a scanned filing, and hand back a structured result with its sources attached. Our guides to investigating an email address and spotting fake candidates show what that collection involves when it is done by hand.

The judgement stays with the investigator: whether a link is strong enough to act on, what a finding means for the case, and what happens to the person or company at the end of it. The agent collects and organises, and the investigator decides.

Five checks before you rely on one

1. Every finding has a source you can open

Citations only help if you open them, because a sourced finding can still be wrong. Check that each page says what the tool says it does.

A 2025 study involving 22 public service media organisations shows why. It found serious sourcing problems, meaning missing, misleading or incorrect attributions, in 31 percent of the answers that four widely used AI assistants gave to news questions. Those were general assistants rather than investigation tools, but the failure is the same one to test for.

Test it directly. Take twenty findings from a few results, open every source, and count how many support the claim as written. Watch for the near misses: a real page that mentions a different person with the same name, or an article that says something weaker than the summary above it.

Sources need dates as well, because pages change and come down, and a finding is only as current as the day its source was read. The Berkeley Protocol on Digital Open Source Investigations, published by the UN human rights office and UC Berkeley's Human Rights Center, asks human investigators to record the search paths that led to each piece of material, when and by whom it was collected, and their method in enough detail that someone else could repeat it. A tool that does the searching for you should keep that record for you.

2. It says "unknown" instead of guessing

For an investigator, a blank field is cheap and a wrong one is expensive. A blank tells you to keep looking. A confident wrong answer can end up in a report.

Language models lean the other way by default. A 2025 paper on why language models hallucinate argues that standard training and evaluation reward a confident guess over an admission of uncertainty. The broadcasters' study above found major accuracy issues, including hallucinated details and outdated information, in 20 percent of the answers it reviewed.

So test for restraint as well as recall. Include a few subjects with thin public footprints and see whether the tool returns gaps or fills them. The result you want is a list of what was established, what was not, and why.

3. It gets the right person among people who share a name

In the 2010 US Census, 2,442,977 people had the surname Smith, and each of the ten most common surnames was shared by more than a million people. A wrong identity match is one of the most expensive mistakes in people research, because everything found after it belongs to someone else.

A good tool anchors to what you already know, such as an employer, a city, an email address or a date of birth, and explains why each new account or record belongs to the same person. Test it with a subject whose name is common, and with one who has a namesake in the same field or city. Read the reasons it gives for each link, not only the links.

4. Collection stays passive

Open-source collection should not touch the subject. The Berkeley Protocol draws its definition of open source information along that line: such information "does not involve interacting with or soliciting information from individual Internet users". For a tool, that rules out logging into the subject's accounts, triggering password resets, messaging people, or anything else that would tell someone they are being looked at or that relies on access nobody authorised. Ask the vendor directly what the tool does to collect, and look for anything in the output that could only have come from an active step.

Ask where your queries go as well. The names and identifiers you search are sensitive in their own right, and you should know how long they are kept and who can see them.

5. It holds up on cases you already know

Your own closed cases make the best test, because you know what the right answer was and how hard it was to find. The next section sets out how to build one.

Run the same case twice as well. Systems built on language models can give different answers to the same question even with settings meant to make them repeatable, and a 2024 study of "deterministic" model settings measured variations in accuracy of up to 15 percent across runs on standard benchmarks. Some variation is normal, and it should sit at the edges of a result, such as an extra old account found one time and not the other. The core identity should not change between runs.

Build your own test set

Twenty to thirty closed cases should surface most of a tool's failure modes. Pick them so that each one tests something specific:

IncludeWhat it tests
Subjects with rich public footprintsRecall, and whether sources are cited properly
Subjects with thin footprintsWhether the tool reports gaps instead of filling them
Common names, and pairs of people who share a nameIdentity resolution
An identifier that changed hands, such as a recycled phone number or a reassigned work email addressWhether the tool checks dates before linking
Names in other scripts or languagesTransliteration and non-English sources
A company with a similarly named counterpartEntity resolution for organisations

Score every field as correct, wrong or missing, and weigh a wrong answer as worse than a missing one. Separately, count the citations that do not support the claim they are attached to.

The first half of that is how RECON scores. RECON is the people research benchmark we built, externally validated by HUD, and it covers 140 real people and 514 verified fields, with roughly a quarter of the people paired with a same-named decoy. Its headline metric is weighted accuracy: correct answers minus wrong answers, divided by every field. A wrong answer costs a point, and a blank costs nothing.

The most accurate system on the benchmark is our highest tier. In the published results it answered 348 of the 514 fields correctly, answered 69 wrongly and left 97 blank, which is roughly one wrong answer in every six it gave. Every system tested got some answers wrong, so a person should read the result before anyone acts on it. The full results and methodology are published.

What Sixtyfour's agent returns

Sixtyfour builds AI research agents for people and company investigations. The agent starts from the identifiers you have, such as a name, email address, phone number or company, researches across public records, breach records, social profiles, code hosting, forums, archived pages, filings and news, and returns findings with a source attached to each one. It says which accounts and records it believes belong to the same person and why, and reports what it could not establish as open rather than filling it in, so an investigator can check every finding before acting on it.

The checks, in brief

  • Open the sources. Count how many support the claim as written, and check they are dated.
  • Give it thin footprints and see whether it admits gaps.
  • Give it common names and pairs of people who share a name, and read its reasons for each link.
  • Ask what it does to collect, and where your queries go.
  • Run your own closed cases, twice, and score wrong answers as worse than blank ones.

Run these checks on your own cases and you will know where a tool's results can be trusted and where a person still needs to look.

“For an investigator, a blank field is cheap and a wrong one is expensive.”
Saarth Shah Co-Founder & CEO, Sixtyfour
1200 × 630 — ready to share

Frequently asked

Research agents that take a starting identifier, such as a name, email address, username, phone number or company, and search public sources to build a sourced profile of the person or organisation behind it, including accounts, records, connections and adverse media. They are used in trust and safety, fraud, due diligence, hiring and open-source intelligence work.

It can do much of the collection: running searches across many sources, pivoting from one identifier to the next, reading documents in other languages and organising what it finds with sources attached. Judging what the evidence means, deciding how confident a link is and deciding what to do about it still need an investigator.

Not for the decisions. AI agents shorten the collection and cross-referencing that take up much of an investigation's time, but every system on our RECON benchmark still got some answers wrong, so a person needs to review the result before anyone acts on it.

Open the sources. Take a sample of findings, open every citation, and check that the page exists and says what the tool claims it says. Watch for real pages about a different person with the same name, which are harder to spot than links that do not work.

RECON is a people research benchmark built by Sixtyfour and externally validated by HUD. It tests how accurately AI agents find verified facts about 140 real people across 514 fields, pairs roughly a quarter of the people with a same-named decoy, and scores a wrong answer as worse than a blank one.

Interact with the person being researched. It should not log into the subject's accounts, trigger password resets, message anyone or take any other step that tells the subject someone is looking or relies on access nobody authorised. The Berkeley Protocol on Digital Open Source Investigations defines open source information along the same line: it does not involve interacting with or soliciting information from individual internet users.

Get started

See how Sixtyfour's research agent works through a case, with a source behind every finding.

Request a Demo
  1. What these tools do well
  2. Five checks before you rely on one
  3. 1. Every finding has a source you can open
  4. 2. It says "unknown" instead of guessing
  5. 3. It gets the right person among people who share a name
  6. 4. Collection stays passive
  7. 5. It holds up on cases you already know
  8. Build your own test set
  9. What Sixtyfour's agent returns
  10. The checks, in brief