How to run an OSINT investigation

One case from the first question to the written report: how to scope it, pivot between identifiers, weigh each link, preserve the evidence and say how sure you are.

OSINTOpen Source InvestigationTrust & SafetyFraud InvestigationsDue DiligenceIdentity Resolution
How to run an OSINT investigation

To run an OSINT investigation, start with a question you can answer and a written scope, list the identifiers you already hold, and pivot from each one while logging every search. Weigh each link by how rare the shared value is and whether it is independent of your other links, corroborate and date every finding, and save each source with its address, a timestamped capture and a hash. Stop when the question is answered, and write each finding with a stated likelihood and the evidence behind it. Keep the work passive throughout: read what is public, and never contact the subject.

This guide follows one invented case from the first question to the write-up. Our guides to email addresses, phone numbers, related accounts and website ownership go deeper on the individual pivots.

Start an OSINT investigation with a question you can answer

A marketplace's trust and safety team receives nine non-delivery complaints in ten days about a new seller, Fenwray Optics, which lists camera lenses well below market price. Someone on the team thinks it looks like quillaperture, a lens seller banned in May for the same pattern. The two accounts share no phone number, email address, device or payout account on the platform, so the account records alone cannot answer it.

"Find out everything about Fenwray Optics" has no end point, and it sends the analyst after every person the shop has touched. A workable question names what you want to know and the decision it serves. Here: is Fenwray Optics run by the same operator as quillaperture? If it is, the team treats the shop as ban evasion and holds its payouts.

Write the scope down beside the question. The Berkeley Protocol on Digital Open Source Investigations, first published in 2020 by the UN human rights office and the Human Rights Center at UC Berkeley's law school, asks for an online investigation plan before the searching starts, setting out the objectives, priorities, strategy and a timeline. In this case the scope covers both shops' identifiers. The buyers who complained, the operator's family and anything that would need contact with the seller are out. The time box is one afternoon.

Then list what you hold, and where each item came from:

IdentifierSourceFirst pivot
Shop name and handle, fenwray_opticsSeller profileThe same strings elsewhere
Email, fenwray.optics@example.comAccount record, verified at signupBreach history and attached accounts
Phone, +1 614 555 0172Account record and shop pageAds, listings and directories
Website, fenwrayoptics.exampleShop descriptionRegistration date and archived copies
quillaperture's handle, email and phoneThe banned accountThe same searches, for comparison

Note which items the platform verified. A verified email or phone number was in the operator's hands on the day of signup. A shop name or a founder's name on an About page is whatever the operator chose to type.

Pivot from each identifier, and keep a log

A pivot uses one identifier to find the next. Work through them one at a time and log every step as you go: the time, what you searched, where, what came back and what you did with it, dead ends included. The Berkeley Protocol describes an investigation as a cycle of inquiry, assessment, collection, preservation, verification and analysis that may be repeated many times as new information opens new lines of inquiry, and it asks investigators to document their activities in every phase. The log is what lets a colleague repeat your path, and it keeps the same search from being run three times.

In the case, four pivots matter.

  • The website. Its registration record shows it was created in August 2026, and its earliest copy in a public web archive is from September. Its About page says the business has traded since 2009.
  • The email address. No breach history and no trace before September, which fits a new address.
  • The handle. fenwray_optics appears nowhere else.
  • The phone number. It appears in a classified ad from March 2024 for a used telephoto lens, which gives a contact address of quillaperture@example.net.

That last pivot is the lead. The rest of the method is about how much weight it can bear.

Three questions decide it.

How rare is the shared value? A link is worth more the fewer people could share it by chance. The string quillaperture has been used by one account on the platform, and the 2024 ad is the only other place it turns up. A shared city or a common first name would be worth almost nothing. Our guide to linking related accounts sets out how to weigh a shared value by the number of accounts that have it.

Is it independent of the other links? Two links that come from one fact count once. The shop page and the website list the same phone number because the operator typed it into both. That is a single link.

Is it the same person, or only the same name? The website's About page names a founder, Theo Halvessy. A search finds a Theo Halvessy who works as a wedding photographer in another state, with a portfolio going back years. It is the same unusual name in a neighbouring field, and nothing else connects them: no shared phone number, email address, handle or image, and his own site lists a different number. Operators can put anyone's name on a page, including a real person's. He stays out of the case, and the team collects nothing more about him.

Then try to break each link. Ask what would make it wrong. Here the answer is that the phone number could have passed to someone else between 2024 and 2026. Doing this on purpose matters. A 2026 study of confirmation bias in police investigations cites an earlier experiment in which participants told explicitly to try to falsify their first hypothesis showed less confirmation bias, while participants who only generated alternative hypotheses did not.

Corroborate and date every finding

A finding is corroborated when a second source supports it without having taken the fact from the first. Trace each fact to where it first appeared. The Berkeley Protocol calls this provenance, and suggests describing the earliest version you find as the "first copy found online", because the original may have been removed. Two articles quoting one press release are one source, and so are two profiles the operator filled in themselves.

In the case, the second source comes from the platform's own listings. A quillaperture listing from April includes a photograph of a lens with its serial number visible on the barrel. A Fenwray listing from September shows the same serial number, in a different photograph against a different background. That link does not depend on the phone number at all. A resale could explain it, so it supports the conclusion rather than settling it alone.

Then date every finding twice: when the thing happened, and when you saw it. Identifiers change hands and pages are edited, so an old record may describe someone else. The Berkeley Protocol notes that texts written at the time of the events they describe tend to be treated as more reliable than those produced long after.

Dated, the case becomes a sequence: the ad in March 2024, quillaperture's listings from January to May 2026, the ban in May, the domain registered in August, the first Fenwray listing in September. The phone number appears under no other name anywhere in that period. That does not prove it never changed hands, and the report will say so.

Preserve your sources so others can check the work

Pages disappear. A Pew Research Center analysis found that a quarter of the webpages that existed at some point between 2013 and 2023 were no longer accessible by October 2023, and that 38 percent of pages from 2013 were gone. An old classified ad is exactly the kind of page that expires.

For each page a finding rests on, keep:

  • its URL;
  • a full-page capture that shows the date and time;
  • the page's HTML source;
  • a hash of every saved file, using an algorithm from the Secure Hash Standard such as SHA-256, so anyone can later confirm the file has not changed;
  • a log entry saying who captured it and when.

That list follows the collection guidance in the Berkeley Protocol. A screenshot on its own is weak evidence. As a guide to archiving from the investigative newsroom Bellingcat puts it, "Screenshots can be easily forged." A public archive copy helps. The Internet Archive's Save Page Now gives a permanent URL for a single page. Public copies are open to anyone with the address, though, so keep your own copy as well.

When to stop

Stop when the question has an answer at the confidence the decision needs. The leads will not run out on their own. Three signs tell you that you are there: the next pivots return what you already have, the next pivot leads outside the scope, or the time box is spent.

In the case, the team has two independent links and a dated sequence. The remaining leads point at the photographer and at the operator's real-world identity. Neither is needed to answer the question, and following them would mostly mean collecting information about people outside it. The Berkeley Protocol's principle of data minimisation says online content should only be collected if it is relevant to a particular investigation. So the team stops, and records what it did not establish as open.

Write the findings with stated confidence

Readers take the same words to mean very different odds. In a 1964 essay published by the CIA, Sherman Kent recalled a 1951 intelligence estimate that called an attack on Yugoslavia that year a "serious possibility". Kent meant odds of about 65 to 35 in favour of an attack. The State Department officials reading it had taken it to mean much lower odds, and when Kent asked the colleagues who had agreed the wording, their answers ran from about 20 to 80 up to 80 to 20.

A fixed scale is meant to close that gap. The US intelligence community's analytic standards, ICD 203, tie each likelihood word to a range, so that "likely" means 55 to 80 percent and "very likely" means 80 to 95 percent. They also ask analysts to keep the likelihood of a judgement separate from their confidence in its basis, to separate the underlying information from assumptions and judgements, and to name what would change the judgement. They were written for intelligence analysts, and the habits suit any team writing for someone who has to decide.

The case's findings, written that way:

  • Fenwray Optics is very likely (80 to 95 percent) run by the same operator as quillaperture. This rests on two independent links: the shop's phone number appears in a March 2024 ad that uses quillaperture's handle, and listings from both shops show a lens with the same serial number. Our confidence is moderate, because the phone link assumes the number did not change hands between 2024 and 2026. A record of another holder in that period would lower the likelihood.
  • The website's claim to have traded since 2009 is not supported by any record found. The website was registered in August 2026, and its earliest archived copy is from September 2026.
  • Open: who the operator is. The About page names a founder, and the one person found with that name has no connection to either shop in any record reviewed.

Every sentence points to a logged and preserved source, so a reviewer can check each one before the team acts.

Working passively and fairly

Open-source research reads what is already public. The Berkeley Protocol draws the line at contact: information acquired from internet users by communicating with them counts as closed source. In practice that means no messages to the seller, no follow requests, no password reset forms and no test logins. Contact also tells the operator someone is looking, which can lead them to delete the accounts you were about to find.

Do not deceive people to get information. The Protocol warns that misrepresentation can damage an investigation's credibility and contaminate the information collected.

Watch your own footprint. Some platforms tell people who looked at them: depending on the viewer's privacy settings, LinkedIn shows a profile's owner the viewer's name, headline, location and industry. Check those settings before you start.

Collect as little about other people as the question allows. Every investigation passes bystanders: the photographer who shares a name, the buyers, the previous holder of a phone number. The Berkeley Protocol lists bystanders among the people an investigation may need to protect. Leave them out of the record unless the question needs them.

Reading the results fairly

A finding is only as good as the record behind it, and the record is uneven.

Quiet is normal. New identifiers, careful people and people who rarely post all leave thin trails. When a search comes back empty, try another identifier before reading anything into the silence.

Identifiers change hands. Phone numbers are recycled, domains lapse and are registered again, and accounts are sold. Dates are what separate one holder from the next.

Names are claims. Anyone can put a name on a page, so a name match needs an identifier match behind it.

Some people are harder to find than others. The Berkeley Protocol notes that online information may be unevenly available from certain groups or segments of society, so a thin result may describe the record more than the person.

And "very likely" leaves a real chance of being wrong. That is why a person reviews the evidence before anyone acts on it.

Where Sixtyfour fits

Much of this method is collection: running each identifier across many sources, following it to the next one, and keeping a source for every fact. Sixtyfour's agent does that part. It starts from an identifier, such as an email address, phone number, username, name or company, researches public sources, links the findings that belong to the same person or company, and returns every finding with its source. Anything it could not establish is reported as open. The results go to your team, and a person decides what happens next. There is more on how investigation teams use it on our investigations page.

The short version

Write down a question you can answer and its scope, list the identifiers you hold, and pivot from each one with a log. Count a link only when the shared value is rare and independent of your other links, try to break it, then corroborate and date every finding and preserve each source with a capture and a hash. Stop when the question is answered, and write each finding with a stated likelihood, the evidence behind it, and what is still open.

“Stop when the question has an answer at the confidence the decision needs.”
Saarth Shah Co-Founder & CEO, Sixtyfour
1200 × 630 — ready to share

Frequently asked

Write down a question you can answer and its scope, list the identifiers you hold, and pivot from each one while logging every search. Weigh each link by how rare the shared value is, corroborate findings from independent sources, date them, and preserve the pages they came from. Stop when the question is answered, and write each finding with a stated likelihood and the evidence behind it.

Pivoting is using one identifier to find the next: an email address leads to a username, the username to an old forum account, and that account to a phone number. Each pivot needs a check that the new record belongs to the same person, because names, usernames and phone numbers are shared and reused. Logging every pivot lets a colleague follow the same path and check it.

Passive research reads what is already public without touching the subject: no messages, no follow requests, no password reset forms and no test logins. It makes it less likely that the subject learns someone is looking, and it keeps the evidence uncontaminated. The Berkeley Protocol on Digital Open Source Investigations treats information obtained by communicating with internet users as closed source rather than open source.

Find the same fact in a second source that did not get it from the first, check when each source was created, and test the link by looking for evidence that would break it. Two pages repeating one claim, or one person typing the same phone number into two profiles, count as a single source.

For each page you rely on, record its URL, a full-page capture showing the date and time, and the page's source code, and compute a hash of each saved file so any later change can be detected. A copy in a public web archive helps too, but public copies are open to anyone with the address, so keep a private copy as well.

Use a fixed set of likelihood words, say what range each one means, and give the evidence and assumptions behind every judgement. The US intelligence community's analytic standards, for example, tie "likely" to 55 to 80 percent and "very likely" to 80 to 95 percent, and keep statements of likelihood separate from statements of confidence.

Get started

See how Sixtyfour researches a case from a single identifier, with a source behind every finding.

Request a Demo
  1. Start an OSINT investigation with a question you can answer
  2. Pivot from each identifier, and keep a log
  3. How strong is a link?
  4. Corroborate and date every finding
  5. Preserve your sources so others can check the work
  6. When to stop
  7. Write the findings with stated confidence
  8. Working passively and fairly
  9. Reading the results fairly
  10. Where Sixtyfour fits
  11. The short version