RECON

A benchmark measuring how well AI agents uncover and verify hard-to-find facts about people.

140
people
514
verified fields
98.5%
human agreement
Built by Sixtyfour
Externally validated by HUD

A Benchmark for People Intelligence

Starting with a name and a few identifying details, agents must recover specific facts without confusing people who share a name.

Deep research benchmarks like BrowseComp, DeepSearchQA, and SealQA measure how well systems synthesize general information across the web. Finding and verifying facts about one specific person is a different problem: people research.

RECON is the first benchmark dedicated to the people research needed for due diligence, fraud investigations, and data enrichment. It tests a system's ability to find and verify datapoints about real individuals from sources that are scattered, unindexed, and often conflicting: archived school records, old forum posts, and public records. The dataset covers 140 people and 514 verified fields, spanning direct lookups, multi-step investigations, and same-name disambiguation.

Figure 1 Weighted accuracy
02040Sixtyfour High54.3%Sixtyfour Medium44.7%Parallel Ultra 2x41.1%Parallel Ultra 8x37.5%Grok 4.20-ma36.2%Grok 4.636.0%Grok 4.330.9%Sixtyfour Low28.8%Parallel Ultra27.2%Exa agent xhigh23.9%GPT-5.6-sol xhigh20.4%Gemini 3.1 Pro13.4%DeepSeek V4 Pro9.3%Claude Haiku 4.51.6%Weighted accuracy / %02040Sixtyfour High54.3%Sixtyfour Medium44.7%Parallel Ultra 2x41.1%Parallel Ultra 8x37.5%Grok 4.20-ma36.2%Grok 4.636.0%Grok 4.330.9%Sixtyfour Low28.8%Parallel Ultra27.2%Exa agent xhigh23.9%GPT-5.6-sol xhigh20.4%Gemini 3.1 Pro13.4%DeepSeek V4 Pro9.3%Claude Haiku 4.51.6%Weighted accuracy / %

Results on 140 people and 514 fields. Weighted accuracy = (correct − incorrect) / all fields. Sixtyfour systems are black; external systems are grey.

Figure 2 Precision and recall
0255075100Sixtyfour High83.5%Sixtyfour Medium83.2%Parallel Ultra 2x80.4%Parallel Ultra 8x78.1%Grok 4.20-ma77.0%Grok 4.677.4%Grok 4.376.2%Sixtyfour Low70.3%Parallel Ultra72.9%Exa agent xhigh73.2%GPT-5.6-sol xhigh74.2%Gemini 3.1 Pro70.4%DeepSeek V4 Pro79.3%Claude Haiku 4.555.7%Recall (square) to precision (circle) / %0255075100Sixtyfour High83.5%Sixtyfour Medium83.2%Parallel Ultra 2x80.4%Parallel Ultra 8x78.1%Grok 4.20-ma77.0%Grok 4.677.4%Grok 4.376.2%Sixtyfour Low70.3%Parallel Ultra72.9%Exa agent xhigh73.2%GPT-5.6-sol xhigh74.2%Gemini 3.1 Pro70.4%DeepSeek V4 Pro79.3%Claude Haiku 4.555.7%Recall (square) to precision (circle) / %

Precision = correct / nonblank answers. Recall = correct / all 514 fields. DeepSeek counts are uniquely inferred from its rounded weighted accuracy and precision.

Results across 514 fields, ordered by weighted accuracy.
System Weighted accuracy Precision Recall Correct / Wrong / Missing Historical p50
Sixtyfour High 54.3% 83.5% 67.7% 348 / 69 / 97 459s
Sixtyfour Medium 44.7% 83.2% 56.0% 288 / 58 / 168 223s
Parallel Ultra 2x 41.1% 80.4% 54.3% 279 / 68 / 167 834s
Parallel Ultra 8x 37.5% 78.1% 52.1% 268 / 75 / 171 678s
Grok 4.20-ma 36.2% 77.0% 51.6% 265 / 79 / 170
Grok 4.6 36.0% 77.4% 50.8% 261 / 76 / 177
Grok 4.3 30.9% 76.2% 44.9% 231 / 72 / 211 15s
Sixtyfour Low 28.8% 70.3% 49.8% 256 / 108 / 150 230s
Parallel Ultra 27.2% 72.9% 43.4% 223 / 83 / 208 589s
Exa agent xhigh 23.9% 73.2% 37.7% 194 / 71 / 249
GPT-5.6-sol xhigh 20.4% 74.2% 31.3% 161 / 56 / 297
Gemini 3.1 Pro (high) 13.4% 70.4% 23.2% 119 / 50 / 345 87s
DeepSeek V4 Pro (high) 9.3% 79.3% 12.6% 65 / 17 / 432*
Claude Haiku 4.5 1.6% 55.7% 7.6% 39 / 31 / 444

Historical p50: June 16–17, 2026, measured separately. *Counts reconstructed from reported weighted accuracy and precision.

What makes people research hard

People research is challenging in three ways.

Identity resolution

The internet is full of people named "Saarth Shah." Finding facts about the right one requires anchoring to known ground truth and verifying each new connection before extending further. A wrong identity match produces a confidently wrong answer, and every finding built on it inherits the error.

Causal chains

Many people research tasks can't be answered with a single search. Confirming whether someone held a prior directorship at a dissolved company might require: finding their current business registration, tracing a shared address to a dissolved entity, then searching that entity's filings to surface their prior role. Each step depends on the previous one.

Fragmented signals

People leave traces across dozens of platforms. Unlike corporate filings, this information is scattered, often unindexed, and rarely conclusive on its own. A key fact might be a single line buried in a scanned PDF, public records, an old forum post, or require reconciling conflicting signals across multiple sources.

How we built the benchmark

People and fields

We compiled 140 individuals across industries, roles, and digital footprint levels, avoiding public figures whose information is too widely indexed to stress-test the benchmark. The dataset covers 514 verified facts that reflect what comes up most in real people research tasks. Every field was verified against a primary source with an unambiguous identity link to the person. Fields relying on inference were rejected.

The people span more than 60 occupational and geographic strata across more than a dozen countries, with digital footprints ranging from heavily indexed to near-invisible. Roughly a quarter are paired with a same-named decoy. The fields cover professional history, education, contact and identity records, and online accounts. A name or username match alone was not sufficient to establish a reference answer.

A sample task

Each entry pairs a person with a set of fields to find. person_info gives the agent its ground-truth anchor, and struct defines what it needs to return. This example shows one field from Zack Kline's task, formatted for readability:

Task
{
"person_info": "Zack Kline, founder of A.I.R. Lawn Care in Rockville, Maryland",
"struct": {
"company_original_tagline": "The company's original tagline"
}
}

company_original_tagline requires following a historical lead. In a selected Sixtyfour run, local coverage identified the company's former domain. A 2012 archive of its homepage contained “Join us, and take the A.I.R. Dare ©.” The reference answer is shown separately below and is not supplied to the agent:

Expected
{
"company_original_tagline": "Take the A.I.R. Dare ©"
}

Each field is graded by a GPT-4.1-mini judge as correct, incorrect, or missing, ignoring format differences. A 200-field audit found 98.5% agreement with human review. Weighted accuracy is computed as (correct − incorrect) / total evaluated fields; standard accuracy is simply correct / total evaluated fields. Precision is correct divided by nonblank answers, and recall is correct divided by all fields. Missing answers remain in the denominator.

External systems run through the public benchmark scripts, with web search and code execution where supported. Results describe the evaluated system and its tools, rather than the underlying model in isolation.

What the numbers tell us

Three patterns emerge clearly across the benchmark. We walk through each one in turn.

1. Specialized research agents lead

Sixtyfour High leads at 54.3% weighted accuracy, followed by Sixtyfour Medium at 44.7%. Parallel Ultra 2x is the highest-scoring external configuration at 41.1%. These results show the performance of complete research systems. They do not isolate the contribution of the model, tool access, or research strategy.

2. Both incorrect and missing answers matter

A wrong answer is harder to catch than a missing one, and in many contexts more damaging. We report weighted accuracy — (correct − incorrect) / total evaluated fields — as our primary metric to account for this. The breakdown below shows correct, incorrect, and missing fields per provider.

Figure 3 Weighted accuracy breakdown
0255075100Sixtyfour HighSixtyfour MediumParallel Ultra 2xParallel Ultra 8xGrok 4.20-maGrok 4.6Grok 4.3Sixtyfour LowParallel UltraExa agent xhighGPT-5.6-sol xhighGemini 3.1 ProDeepSeek V4 ProClaude Haiku 4.5CorrectIncorrectWeighted accuracy0255075100Sixtyfour HighSixtyfour MediumParallel Ultra 2xParallel Ultra 8xGrok 4.20-maGrok 4.6Grok 4.3Sixtyfour LowParallel UltraExa agent xhighGPT-5.6-sol xhighGemini 3.1 ProDeepSeek V4 ProClaude Haiku 4.5CorrectIncorrectWeighted accuracy

Correct and incorrect answers are shown as shares of all 514 fields; the unfilled remainder is missing. The dotted portion shows weighted accuracy.

3. People research is an unsolved problem

Sixtyfour High returned 348 correct answers, 69 incorrect answers, and 97 missing answers. Its 83.5% precision means 16.5% of supplied answers were wrong. Even the highest-scoring system leaves many facts unresolved. In high-stakes contexts, critical findings should still be verified against primary sources.

What this benchmark does not say

RECON measures accuracy on a fixed set of fields; different field selection could produce different rankings. Single-run results can vary with source availability, provider behavior, and configuration. They do not measure the quality of an open-ended research report. Refusals, tool failures, and unsuccessful searches can all reduce answer coverage and should be distinguished when interpreting traces.

Historical p50 latencies were measured separately on June 16–17, 2026. Configurations may have changed since then, so these timings should not be interpreted as current performance.

Run the benchmark

The benchmark repository includes runner scripts and evaluation instructions. The dataset is available by request because it contains personal information.

For people research beyond the benchmark

Sixtyfour Research RECON Benchmark · Sixtyfour