A Benchmark for People Intelligence
Starting with a name and a few identifying details, agents must recover specific facts without confusing people who share a name.
Deep research benchmarks like BrowseComp, DeepSearchQA, and SealQA measure how well systems synthesize general information across the web. Finding and verifying facts about one specific person is a different problem: people research.
RECON is the first benchmark dedicated to the people research needed for due diligence, fraud investigations, and data enrichment. It tests a system's ability to find and verify datapoints about real individuals from sources that are scattered, unindexed, and often conflicting: archived school records, old forum posts, and public records. The dataset covers 140 people and 514 verified fields, spanning direct lookups, multi-step investigations, and same-name disambiguation.
Results on 140 people and 514 fields. Weighted accuracy = (correct − incorrect) / all fields. Sixtyfour systems are black; external systems are grey.
Precision = correct / nonblank answers. Recall = correct / all 514 fields. DeepSeek counts are uniquely inferred from its rounded weighted accuracy and precision.
| System | Weighted accuracy | Precision | Recall | Correct / Wrong / Missing | Historical p50 |
|---|---|---|---|---|---|
| Sixtyfour High | 54.3% | 83.5% | 67.7% | 348 / 69 / 97 | 459s |
| Sixtyfour Medium | 44.7% | 83.2% | 56.0% | 288 / 58 / 168 | 223s |
| Parallel Ultra 2x | 41.1% | 80.4% | 54.3% | 279 / 68 / 167 | 834s |
| Parallel Ultra 8x | 37.5% | 78.1% | 52.1% | 268 / 75 / 171 | 678s |
| Grok 4.20-ma | 36.2% | 77.0% | 51.6% | 265 / 79 / 170 | — |
| Grok 4.6 | 36.0% | 77.4% | 50.8% | 261 / 76 / 177 | — |
| Grok 4.3 | 30.9% | 76.2% | 44.9% | 231 / 72 / 211 | 15s |
| Sixtyfour Low | 28.8% | 70.3% | 49.8% | 256 / 108 / 150 | 230s |
| Parallel Ultra | 27.2% | 72.9% | 43.4% | 223 / 83 / 208 | 589s |
| Exa agent xhigh | 23.9% | 73.2% | 37.7% | 194 / 71 / 249 | — |
| GPT-5.6-sol xhigh | 20.4% | 74.2% | 31.3% | 161 / 56 / 297 | — |
| Gemini 3.1 Pro (high) | 13.4% | 70.4% | 23.2% | 119 / 50 / 345 | 87s |
| DeepSeek V4 Pro (high) | 9.3% | 79.3% | 12.6% | 65 / 17 / 432* | — |
| Claude Haiku 4.5 | 1.6% | 55.7% | 7.6% | 39 / 31 / 444 | — |
Historical p50: June 16–17, 2026, measured separately. *Counts reconstructed from reported weighted accuracy and precision.
What makes people research hard
People research is challenging in three ways.
Identity resolution
The internet is full of people named "Saarth Shah." Finding facts about the right one requires anchoring to known ground truth and verifying each new connection before extending further. A wrong identity match produces a confidently wrong answer, and every finding built on it inherits the error.
Causal chains
Many people research tasks can't be answered with a single search. Confirming whether someone held a prior directorship at a dissolved company might require: finding their current business registration, tracing a shared address to a dissolved entity, then searching that entity's filings to surface their prior role. Each step depends on the previous one.
Fragmented signals
People leave traces across dozens of platforms. Unlike corporate filings, this information is scattered, often unindexed, and rarely conclusive on its own. A key fact might be a single line buried in a scanned PDF, public records, an old forum post, or require reconciling conflicting signals across multiple sources.
How we built the benchmark
People and fields
We compiled 140 individuals across industries, roles, and digital footprint levels, avoiding public figures whose information is too widely indexed to stress-test the benchmark. The dataset covers 514 verified facts that reflect what comes up most in real people research tasks. Every field was verified against a primary source with an unambiguous identity link to the person. Fields relying on inference were rejected.
The people span more than 60 occupational and geographic strata across more than a dozen countries, with digital footprints ranging from heavily indexed to near-invisible. Roughly a quarter are paired with a same-named decoy. The fields cover professional history, education, contact and identity records, and online accounts. A name or username match alone was not sufficient to establish a reference answer.
A sample task
Each entry pairs a person with a set of fields to find. person_info gives
the agent its ground-truth anchor, and struct defines what it needs
to return. This example shows one field from Zack Kline's task, formatted
for readability:
{"person_info": "Zack Kline, founder of A.I.R. Lawn Care in Rockville, Maryland","struct": {"company_original_tagline": "The company's original tagline"}}
company_original_tagline requires following a historical lead.
In a selected Sixtyfour run, local coverage identified the company's former
domain. A 2012
archive of its homepage contained “Join us, and take the A.I.R. Dare ©.”
The reference answer is shown separately below and is not supplied to the agent:
{"company_original_tagline": "Take the A.I.R. Dare ©"}
Each field is graded by a GPT-4.1-mini judge as correct, incorrect, or missing,
ignoring format differences. A 200-field audit found
98.5% agreement with human review. Weighted accuracy is computed as
(correct − incorrect) / total evaluated fields; standard accuracy is
simply correct / total evaluated fields. Precision is correct
divided by nonblank answers, and recall is correct divided by all fields.
Missing answers remain in the denominator.
External systems run through the public benchmark scripts, with web search and code execution where supported. Results describe the evaluated system and its tools, rather than the underlying model in isolation.
What the numbers tell us
Three patterns emerge clearly across the benchmark. We walk through each one in turn.
1. Specialized research agents lead
Sixtyfour High leads at 54.3% weighted accuracy, followed by Sixtyfour Medium at 44.7%. Parallel Ultra 2x is the highest-scoring external configuration at 41.1%. These results show the performance of complete research systems. They do not isolate the contribution of the model, tool access, or research strategy.
2. Both incorrect and missing answers matter
A wrong answer is harder to catch than a missing one, and in many contexts more
damaging. We report weighted accuracy — (correct − incorrect) / total
evaluated fields — as our primary metric to account for this. The breakdown
below shows correct, incorrect, and missing fields per provider.
Correct and incorrect answers are shown as shares of all 514 fields; the unfilled remainder is missing. The dotted portion shows weighted accuracy.
3. People research is an unsolved problem
Sixtyfour High returned 348 correct answers, 69 incorrect answers, and 97 missing answers. Its 83.5% precision means 16.5% of supplied answers were wrong. Even the highest-scoring system leaves many facts unresolved. In high-stakes contexts, critical findings should still be verified against primary sources.
What this benchmark does not say
RECON measures accuracy on a fixed set of fields; different field selection could produce different rankings. Single-run results can vary with source availability, provider behavior, and configuration. They do not measure the quality of an open-ended research report. Refusals, tool failures, and unsuccessful searches can all reduce answer coverage and should be distinguished when interpreting traces.
Historical p50 latencies were measured separately on June 16–17, 2026. Configurations may have changed since then, so these timings should not be interpreted as current performance.
Run the benchmark
The benchmark repository includes runner scripts and evaluation instructions. The dataset is available by request because it contains personal information.
For people research beyond the benchmark