What Is OSINT (Open-Source Intelligence)?

What Is OSINT (Open-Source Intelligence)?

In This Article

  1. The definition, and the word doing the work
  2. Where open-source material comes from
  3. How a finished product actually gets made
  4. Grading the source and grading the claim
  5. The provenance discipline
  6. What makes OSINT hard right now
  7. Where machines help, and where we draw the line

Key Takeaways

The definition, and the word doing the work

Open-source intelligence is intelligence produced from information that can be lawfully obtained without clandestine collection. The word carrying the weight is intelligence: the output is an assessment answering a question someone asked. A folder of links is not OSINT, any more than a stack of photographs is imagery intelligence.

Two official definitions are worth holding side by side.

The added word is commercially. The 2006 language describes material that is public; the 2024 language also covers material that is purchasable — commercial satellite imagery, licensed datasets, brokered records. Same discipline, far larger aperture, harder legal questions. Under that strategy the Director of the CIA serves as the Intelligence Community's OSINT Functional Manager.

The input term has its own definition. The 2012 edition of Army Techniques Publication ATP 2-22.9, Open-Source Intelligence, defines publicly available information as "data, facts, instructions, or other material published or broadcast for general public consumption; available on request to a member of the general public; lawfully seen or heard by any casual observer; or made available at a meeting open to the general public." It also notes a distinction that trips people up: "open-source information can be publicly available but not all publicly available information is open source."

Where open-source material comes from

ATP 2-22.9 sorts open-source media into four families — public speaking, public documents, public broadcasts, and internet web sites — organized by delivery channel rather than by subject. It is old enough to list shortwave radio and business cards, which is why it has aged well: channels change, categories do not.

A working list an analyst would use today, ordered roughly by strength of provenance:

For all of them, the question that matters is not "is it public?" It is: who created this record, when, and what process produced it? That is the difference between a source and a rumor with a URL.

How a finished product actually gets made

OSINT runs through the same process as any other intelligence discipline. Joint Publication 2-0, Joint Intelligence (22 October 2013), describes six interrelated categories of intelligence operations: planning and direction; collection; processing and exploitation; analysis and production; dissemination and integration; and evaluation and feedback. JP 2-0 is explicit that these are not an assembly line — operations often occur nearly simultaneously, or a step is bypassed.

Made concrete with an ordinary question — has activity at this port changed over the past six months?planning and direction converts it into answerable elements of information, because "has activity changed" is not yet answerable and "how many vessel calls were publicly reported per month" is. Collection gathers the material. Processing and exploitation translates, transcribes, extracts, deduplicates, and normalizes it into a form comparable across sources. Analysis and production is the reasoning and the writing. Dissemination and integration puts it in front of the person who has to act. Evaluation and feedback asks whether it answered the question.

The step tooling-first efforts underestimate is processing and exploitation. Collection is the easy part and gets easier every year. Turning a thousand heterogeneous records — different languages, formats, date conventions, and reliability — into something an analyst can compare line by line is where most of the labor sits. In cyber threat intelligence the same burden lands on a platform that normalizes, deduplicates, scores, and enriches every incoming feed before an analyst sees any of it. It is also where the choice between retrieving passages and extracting structured fields starts to matter; our guide on RAG versus extraction covers that trade-off.

Grading the source and grading the claim

A habit worth borrowing far outside government: grade the source and the claim separately. ATP 2-22.9 uses two scales. Source reliability runs from A ("no doubt of authenticity, trustworthiness, or competency") down to F ("no basis exists for evaluating the reliability of the source"). Content credibility runs from 1 ("confirmed by other independent sources; logical in itself") to 8 (cannot be judged).

Two details repay attention. A first-time source is rated F and a first-time claim is rated 8 — which, the publication is careful to say, does not mean unreliable or false; it means there is no basis yet. And the scale distinguishes 6, misinformation ("unintentionally false") from 7, deception ("deliberately false"). Intent is a separate finding from accuracy, and conflating them is a common analytic error.

For the language of the conclusion itself, Intelligence Community Directive 203, Analytic Standards (signed January 2, 2015) sets five analytic standards and nine tradecraft standards. Three matter to anyone writing from open sources. Products must properly describe "quality and credibility of underlying sources, data, and methodologies." They must properly distinguish "between underlying intelligence information and analysts' assumptions and judgments" — applied honestly, the fastest way to improve any assessment. And they must express uncertainty in a fixed vocabulary with published probability bands: almost no chance at 01–05 percent, roughly even chance at 45–55 percent, almost certain at 95–99 percent. ICD 203 also forbids combining a confidence level and a degree of likelihood in one sentence, since "high confidence that it is likely" leaves a reader unsure which is claimed.

The provenance discipline

This is what separates OSINT from careful browsing, and it is the least glamorous part of the discipline. ICD 206, Sourcing Requirements for Disseminated Analytic Products, spells out what must accompany a finished assessment so a reader can judge the sources underneath it. Four mechanisms carry that load:

That last requirement is the one to internalize. Web pages are edited and deleted; accounts are removed; archives are incomplete. If the record is not captured at the moment it is used, the citation decays into an assertion — and the assessment becomes unauditable at exactly the moment someone challenges it.

The same instinct produced the Berkeley Protocol on Digital Open Source Investigations, a joint publication of the UN Office of the High Commissioner for Human Rights and the Human Rights Center at the University of California, Berkeley, School of Law, issued as a United Nations publication in 2022. It sets out methodologies for gathering, analysing, and preserving digital open-source information to a standard that holds up in a legal proceeding, and it is free to read.

What makes OSINT hard right now

Purchased data carries obligations public data does not. A Senior Advisory Group Panel convened by the Director of National Intelligence examined the IC's use of commercially available information; its report was declassified in June 2023. The panel found that the IC collects a significant amount of CAI for mission purposes, that the widespread availability of such data poses counterintelligence risks, and that it carries important implications for U.S. person privacy and civil liberties because CAI can reveal sensitive and intimate information about individuals. Buying a dataset does not retire those questions. It relocates them into a contract.

Authenticity cannot be assumed, and cannot yet be reliably detected. NIST's report AI 100-4, Reducing Risks Posed by Synthetic Content, surveys the technical approaches to digital content transparency — tracking provenance data, digital watermarking, metadata recording, and synthetic-content detection — and treats none as solved. The consequence is directional: capture provenance at ingest, while you still have the headers, timestamps, and retrieval context, rather than reconstructing authenticity from the artifact later.

Volume is not the constraint; requirements are. The failure mode in modern open-source work is rarely too little material. It is a thousand plausible documents, no stated question, and no record of which ones the conclusion rests on.

Where machines help, and where we draw the line

A note on scope, since this is our work. Precision Federal (Precision Delivery Federal LLC, an Iowa limited liability company) builds small models that read through a body of data and produce a written conclusion with every statement traced back to the exact record it came from — a processing-and-exploitation capability with a citation layer. When the deliverable has to run inside an accreditation boundary rather than call a commercial API, the constraints change; we cover that in DoD Impact Levels explained.

Three things we will not do — worth asking of any vendor here. We will not call model output "intelligence." A model can extract, compare, translate, and draft; the analytic judgment and accountability for it belong to a named human being — the same line we hold in cyber threat intelligence work, where attribution calls stay with cleared analysts and the IC. We will not build a system whose citations point only at a live URL. If the record is not captured and hashed at ingest, the citation is decoration and will fail the first time it is tested. We will not collect or process information on U.S. persons outside a customer's own legal authority and written procedures. That is a question for the customer's counsel and oversight officials — not an engineering preference, and not a contractor's call.

Sources: 50 U.S.C. § 3038, statutory note from Pub. L. 109-163 § 931; ODNI/CIA — The IC OSINT Strategy 2024–2026; U.S. Army — ATP 2-22.9, Open-Source Intelligence (10 July 2012); Joint Publication 2-0, Joint Intelligence (22 October 2013); ICD 203 — Analytic Standards; ICD 206 — Sourcing Requirements for Disseminated Analytic Products (both directives are also listed on ODNI's Intelligence Community Directives page); ODNI — Senior Advisory Group Panel Declassified Report on Commercially Available Information; UN OHCHR / UC Berkeley Human Rights Center — Berkeley Protocol on Digital Open Source Investigations; NIST AI 100-4 — Reducing Risks Posed by Synthetic Content. Analysis and framing by Precision AI Academy.

Common questions

What is OSINT? Intelligence produced from information that can be lawfully obtained without clandestine collection. A congressional finding in Section 931 of the FY2006 NDAA, carried as a note to 50 U.S.C. § 3038, defines it as intelligence produced from publicly available information and "collected, exploited, and disseminated in a timely manner to an appropriate audience for the purpose of addressing a specific intelligence requirement." The IC OSINT Strategy 2024–2026 uses a broader phrasing — "publicly or commercially available information."

What counts as publicly available information? ATP 2-22.9 defines it as "data, facts, instructions, or other material published or broadcast for general public consumption; available on request to a member of the general public; lawfully seen or heard by any casual observer; or made available at a meeting open to the general public." The same publication notes that open-source information can be publicly available, but not all publicly available information is open source.

How is OSINT different from just searching the internet? A search returns material. OSINT is a process that starts from a stated requirement and ends in a written assessment whose sourcing a reader can audit. ICD 206 requires source reference citations and source descriptors in covered analytic products, strongly encourages source summary statements, and requires that a record of a dynamic source such as an internet posting be preserved for at least one year.

Can AI models do OSINT? They are useful for processing and exploitation — translation, transcription, deduplication, extraction, and drafting with citations back to specific records. They do not supply analytic judgment or accountability for it, which belong to a named human analyst. NIST AI 100-4 also documents that detecting synthetic content remains unsolved, which is why serious pipelines capture provenance at the moment of collection rather than inferring authenticity afterward.

About Precision AI Academy

Precision AI Academy publishes practical AI news, plain-language analysis, and free courses for builders and working professionals. It is a sister site of Precision Federal, a federal software and AI firm. We verify the numbers, cite the primary sources, and skip the hype.

Need this built?

If you are reading this because it is a live problem rather than a curiosity: this is what Precision Federal — a federal software and AI firm, and the sister company of this site — builds. Specifically, analytic systems over open-source data where every derived conclusion stays traceable to the exact record it came from, so a reviewer can check the chain rather than trust it.

How it usually starts. A short, scoped assessment against your real system and constraints, ending in a written recommendation you keep whether or not you go further. No retainer to have the first conversation.

What we will not do. We do not collect intelligence and we do not acquire data on your behalf. We build the engineering that turns data you already hold into something defensible.

See the threat intelligence capability → Talk to Precision Federal