In This Article
Key Takeaways
- OSINT is a discipline, not a toolkit. The official definitions both end the same way — the work has to answer a specific stated requirement, or it is just a collection.
- The 2006 statutory-note language says "publicly available information." The IC's 2024 strategy says "publicly or commercially available information." That one added word is where most of today's legal and privacy complexity lives.
- Source reliability and claim credibility are graded separately. Army doctrine uses an A–F scale for the source and a 1–8 scale for the content; a reliable outlet can carry an unconfirmed claim.
- ICD 206 requires more than a hyperlink: numbered source citations, source descriptors, and — for a dynamic source such as an internet posting — a preserved record kept for at least a year. A source summary statement is strongly encouraged rather than required.
The definition, and the word doing the work
Open-source intelligence is intelligence produced from information that can be lawfully obtained without clandestine collection. The word carrying the weight is intelligence: the output is an assessment answering a question someone asked. A folder of links is not OSINT, any more than a stack of photographs is imagery intelligence.
Two official definitions are worth holding side by side.
- The 2006 congressional finding. Section 931 of the FY2006 National Defense Authorization Act (Public Law 109-163), carried as a statutory note to 50 U.S.C. § 3038, states that OSINT "is intelligence that is produced from publicly available information and is collected, exploited, and disseminated in a timely manner to an appropriate audience for the purpose of addressing a specific intelligence requirement."
- The 2024 strategy. The IC OSINT Strategy 2024–2026, released by ODNI and CIA on March 8, 2024 and subtitled The INT of First Resort, defines OSINT as "intelligence derived exclusively from publicly or commercially available information that addresses specific intelligence priorities, requirements, or gaps."
The added word is commercially. The 2006 language describes material that is public; the 2024 language also covers material that is purchasable — commercial satellite imagery, licensed datasets, brokered records. Same discipline, far larger aperture, harder legal questions. Under that strategy the Director of the CIA serves as the Intelligence Community's OSINT Functional Manager.
The input term has its own definition. The 2012 edition of Army Techniques Publication ATP 2-22.9, Open-Source Intelligence, defines publicly available information as "data, facts, instructions, or other material published or broadcast for general public consumption; available on request to a member of the general public; lawfully seen or heard by any casual observer; or made available at a meeting open to the general public." It also notes a distinction that trips people up: "open-source information can be publicly available but not all publicly available information is open source."
Where open-source material comes from
ATP 2-22.9 sorts open-source media into four families — public speaking, public documents, public broadcasts, and internet web sites — organized by delivery channel rather than by subject. It is old enough to list shortwave radio and business cards, which is why it has aged well: channels change, categories do not.
A working list an analyst would use today, ordered roughly by strength of provenance:
- Records of institutional origin. Court filings, regulatory and corporate disclosures, procurement notices, patents, standards documents, peer-reviewed literature, official statistics — dated, versioned, and archived by an institution with a reason to keep them straight.
- News and broadcast media. Reported material with a named publisher and a correction process, in many languages. Strong on timeliness, variable on independence.
- User-generated and social content. Enormous volume, immediate, and the weakest provenance in the set — accounts can be edited, deleted, or manufactured outright.
- Imagery, geospatial, and technical records. Purchased overhead imagery, elevation and mapping layers, domain registrations, certificate transparency logs, routing tables, public code repositories.
- Commercially available information. Licensed or brokered datasets about people, firms, vehicles, shipments, and transactions — legally the most fraught category, discussed below.
For all of them, the question that matters is not "is it public?" It is: who created this record, when, and what process produced it? That is the difference between a source and a rumor with a URL.
How a finished product actually gets made
OSINT runs through the same process as any other intelligence discipline. Joint Publication 2-0, Joint Intelligence (22 October 2013), describes six interrelated categories of intelligence operations: planning and direction; collection; processing and exploitation; analysis and production; dissemination and integration; and evaluation and feedback. JP 2-0 is explicit that these are not an assembly line — operations often occur nearly simultaneously, or a step is bypassed.
Made concrete with an ordinary question — has activity at this port changed over the past six months? — planning and direction converts it into answerable elements of information, because "has activity changed" is not yet answerable and "how many vessel calls were publicly reported per month" is. Collection gathers the material. Processing and exploitation translates, transcribes, extracts, deduplicates, and normalizes it into a form comparable across sources. Analysis and production is the reasoning and the writing. Dissemination and integration puts it in front of the person who has to act. Evaluation and feedback asks whether it answered the question.
The step tooling-first efforts underestimate is processing and exploitation. Collection is the easy part and gets easier every year. Turning a thousand heterogeneous records — different languages, formats, date conventions, and reliability — into something an analyst can compare line by line is where most of the labor sits. In cyber threat intelligence the same burden lands on a platform that normalizes, deduplicates, scores, and enriches every incoming feed before an analyst sees any of it. It is also where the choice between retrieving passages and extracting structured fields starts to matter; our guide on RAG versus extraction covers that trade-off.
Grading the source and grading the claim
A habit worth borrowing far outside government: grade the source and the claim separately. ATP 2-22.9 uses two scales. Source reliability runs from A ("no doubt of authenticity, trustworthiness, or competency") down to F ("no basis exists for evaluating the reliability of the source"). Content credibility runs from 1 ("confirmed by other independent sources; logical in itself") to 8 (cannot be judged).
Two details repay attention. A first-time source is rated F and a first-time claim is rated 8 — which, the publication is careful to say, does not mean unreliable or false; it means there is no basis yet. And the scale distinguishes 6, misinformation ("unintentionally false") from 7, deception ("deliberately false"). Intent is a separate finding from accuracy, and conflating them is a common analytic error.
For the language of the conclusion itself, Intelligence Community Directive 203, Analytic Standards (signed January 2, 2015) sets five analytic standards and nine tradecraft standards. Three matter to anyone writing from open sources. Products must properly describe "quality and credibility of underlying sources, data, and methodologies." They must properly distinguish "between underlying intelligence information and analysts' assumptions and judgments" — applied honestly, the fastest way to improve any assessment. And they must express uncertainty in a fixed vocabulary with published probability bands: almost no chance at 01–05 percent, roughly even chance at 45–55 percent, almost certain at 95–99 percent. ICD 203 also forbids combining a confidence level and a degree of likelihood in one sentence, since "high confidence that it is likely" leaves a reader unsure which is claimed.
The provenance discipline
This is what separates OSINT from careful browsing, and it is the least glamorous part of the discipline. ICD 206, Sourcing Requirements for Disseminated Analytic Products, spells out what must accompany a finished assessment so a reader can judge the sources underneath it. Four mechanisms carry that load:
- Source reference citations — numbered endnotes generated each time a source is cited or a judgment depends on it. Required elements include the originator, an unambiguous identifier that lets someone else retrieve the source, a title, and the date of issuance, publication, or posting — with the pointed instruction that "if date of online posting is unknown, then date of online access."
- Source descriptors — short qualitative notes on a source's quality. ICD 206 explicitly allows analysts to "devise their own source descriptors for sources of publicly available information," a quiet acknowledgement that open-source material does not arrive pre-graded.
- Source summary statements — strongly encouraged rather than mandatory: a holistic paragraph on the source base, covering which sources matter most to the key judgments, which corroborate each other, and which conflict.
- Preservation of dynamic sources — when citing something dynamic, such as an internet posting, "a record of the source shall be preserved for retention… for at least one year."
That last requirement is the one to internalize. Web pages are edited and deleted; accounts are removed; archives are incomplete. If the record is not captured at the moment it is used, the citation decays into an assertion — and the assessment becomes unauditable at exactly the moment someone challenges it.
The same instinct produced the Berkeley Protocol on Digital Open Source Investigations, a joint publication of the UN Office of the High Commissioner for Human Rights and the Human Rights Center at the University of California, Berkeley, School of Law, issued as a United Nations publication in 2022. It sets out methodologies for gathering, analysing, and preserving digital open-source information to a standard that holds up in a legal proceeding, and it is free to read.
What makes OSINT hard right now
Purchased data carries obligations public data does not. A Senior Advisory Group Panel convened by the Director of National Intelligence examined the IC's use of commercially available information; its report was declassified in June 2023. The panel found that the IC collects a significant amount of CAI for mission purposes, that the widespread availability of such data poses counterintelligence risks, and that it carries important implications for U.S. person privacy and civil liberties because CAI can reveal sensitive and intimate information about individuals. Buying a dataset does not retire those questions. It relocates them into a contract.
Authenticity cannot be assumed, and cannot yet be reliably detected. NIST's report AI 100-4, Reducing Risks Posed by Synthetic Content, surveys the technical approaches to digital content transparency — tracking provenance data, digital watermarking, metadata recording, and synthetic-content detection — and treats none as solved. The consequence is directional: capture provenance at ingest, while you still have the headers, timestamps, and retrieval context, rather than reconstructing authenticity from the artifact later.
Volume is not the constraint; requirements are. The failure mode in modern open-source work is rarely too little material. It is a thousand plausible documents, no stated question, and no record of which ones the conclusion rests on.
Where machines help, and where we draw the line
A note on scope, since this is our work. Precision Federal (Precision Delivery Federal LLC, an Iowa limited liability company) builds small models that read through a body of data and produce a written conclusion with every statement traced back to the exact record it came from — a processing-and-exploitation capability with a citation layer. When the deliverable has to run inside an accreditation boundary rather than call a commercial API, the constraints change; we cover that in DoD Impact Levels explained.
Three things we will not do — worth asking of any vendor here. We will not call model output "intelligence." A model can extract, compare, translate, and draft; the analytic judgment and accountability for it belong to a named human being — the same line we hold in cyber threat intelligence work, where attribution calls stay with cleared analysts and the IC. We will not build a system whose citations point only at a live URL. If the record is not captured and hashed at ingest, the citation is decoration and will fail the first time it is tested. We will not collect or process information on U.S. persons outside a customer's own legal authority and written procedures. That is a question for the customer's counsel and oversight officials — not an engineering preference, and not a contractor's call.
Sources: 50 U.S.C. § 3038, statutory note from Pub. L. 109-163 § 931; ODNI/CIA — The IC OSINT Strategy 2024–2026; U.S. Army — ATP 2-22.9, Open-Source Intelligence (10 July 2012); Joint Publication 2-0, Joint Intelligence (22 October 2013); ICD 203 — Analytic Standards; ICD 206 — Sourcing Requirements for Disseminated Analytic Products (both directives are also listed on ODNI's Intelligence Community Directives page); ODNI — Senior Advisory Group Panel Declassified Report on Commercially Available Information; UN OHCHR / UC Berkeley Human Rights Center — Berkeley Protocol on Digital Open Source Investigations; NIST AI 100-4 — Reducing Risks Posed by Synthetic Content. Analysis and framing by Precision AI Academy.
Common questions
What is OSINT? Intelligence produced from information that can be lawfully obtained without clandestine collection. A congressional finding in Section 931 of the FY2006 NDAA, carried as a note to 50 U.S.C. § 3038, defines it as intelligence produced from publicly available information and "collected, exploited, and disseminated in a timely manner to an appropriate audience for the purpose of addressing a specific intelligence requirement." The IC OSINT Strategy 2024–2026 uses a broader phrasing — "publicly or commercially available information."
What counts as publicly available information? ATP 2-22.9 defines it as "data, facts, instructions, or other material published or broadcast for general public consumption; available on request to a member of the general public; lawfully seen or heard by any casual observer; or made available at a meeting open to the general public." The same publication notes that open-source information can be publicly available, but not all publicly available information is open source.
How is OSINT different from just searching the internet? A search returns material. OSINT is a process that starts from a stated requirement and ends in a written assessment whose sourcing a reader can audit. ICD 206 requires source reference citations and source descriptors in covered analytic products, strongly encourages source summary statements, and requires that a record of a dynamic source such as an internet posting be preserved for at least one year.
Can AI models do OSINT? They are useful for processing and exploitation — translation, transcription, deduplication, extraction, and drafting with citations back to specific records. They do not supply analytic judgment or accountability for it, which belong to a named human analyst. NIST AI 100-4 also documents that detecting synthetic content remains unsolved, which is why serious pipelines capture provenance at the moment of collection rather than inferring authenticity afterward.