On Friday 31 January 2025, federal websites started going dark. Pages on gender and health at the Centers for Disease Control and Prevention. The diversity module of the Office of Personnel Management’s FedScope database. A national repository of federal law-enforcement misconduct records at the Department of Justice, decommissioned the week before. The exact count of removed datasets has been contested — a Council on Criminal Justice review put it at roughly 3,000 across 2025 — but the pattern was unmistakable to the researchers watching URLs blink out in real time.

That same afternoon, a room at the Harvard T.H. Chan School of Public Health filled with epidemiologists, data scientists and graduate students running wget and archival scrapers against federal domains. They called themselves the Preserving Public Health Data Collective, and the school’s own account of the session describes the plan: mirror whatever they could reach and push it into Harvard Dataverse before it vanished. Jonathan Gilmour, the Chan School data scientist who helped organise it, later said the preservation work had begun back in November 2024 and was nowhere near finished when the purge hit.

government website archive

What actually vanished, and when

Federal agencies maintain roughly 315,000 datasets covering nearly every measurable part of American life, from highway conditions to childhood vaccination rates — the figure comes from the Data.gov catalogue as of August 2025. In the days after 20 January, thousands of pages linked to those datasets began coming down in response to two White House executive orders: one directing agencies to strip references to gender ideology, the other ending federal diversity, equity and inclusion programmes.

NPR reporters tracking the CDC domain documented the scale early. Their 6 February 2025 report catalogued which health pages went dark, which came back, and which returned with altered wording. The Youth Risk Behavior Surveillance System — the CDC’s long-running measure of adolescent substance use, mental health, community-violence exposure and sexual behaviour — was among the affected datasets.

A federal judge ordered health websites restored on 11 February 2025, and many came back. They came back carrying a notice. The restored HHS pages state that the department is complying under court order, and that the page “does not reflect biological reality.”

Federal statistical products had not previously arrived with editorial commentary from political leadership attached.

The scraping sessions

The Chan School datathon was one node in a much larger, largely improvised network. Volunteers coordinated over Slack channels and shared spreadsheets of URLs. Gilmour put the scale problem plainly to NPR: “These federal websites are gigantic, and result in terabytes of data.” Nobody could say with confidence how much had been captured.

The largest single rescue came from a different corner of the same university. On 6 February 2025 the Harvard Law School Library Innovation Lab announced a 16-terabyte archive of Data.gov: more than 311,000 federal datasets harvested across 2024 and 2025, published with digital signatures so that provenance could be checked later.

Other material was caught by the Internet Archive’s End of Term crawl, which has taken snapshots of federal websites at every presidential transition since 2008. More was preserved by the Data Rescue Project, a coalition of data librarians and statisticians launched in February 2025 and described by two ICPSR researchers writing in The Conversation. By August 2025 the project’s portal held roughly 1,100 datasets.

The scraping was not neutral archival work. It was triage.

What the changes look like at the field level

Denice Ross served as U.S. Chief Data Scientist until December 2024 and now tracks federal data at the Federation of American Scientists. Speaking to Marketplace in July 2025, she described a “targeted, surgical removal of data sets” and said her team had identified more than 400 changes to federal forms and surveys since 20 January, most tied to erasing gender-identity and DEI fields.

OPM’s FedScope database stopped publishing racial and ethnic breakdowns of the federal workforce. Pew Research analyst Drew DeSilver, who had used the tool in September 2024 to work out what share of federal employees were Black or Latino, found the diversity module simply gone when he went back to it.

NOAA retired its Billion-Dollar Weather and Climate Disasters database, which had tracked the frequency and cost of major U.S. climate events since 1980. The agency notice, dated 8 May 2025, said the product would receive no updates beyond calendar year 2024. That one has a coda: in October 2025 the nonprofit Climate Central revived the database and hired Adam Smith, the NOAA scientist who had run it.

The CDC’s National Notifiable Diseases Surveillance System dropped questions on sexual orientation and gender identity, which Ross said would make it harder to see how disease burden falls differently across those populations. The Bureau of Justice Statistics removed gender-identity questions from upcoming rounds of the National Crime Victimization Survey — flagged by Steve Pierson, director of science policy at the American Statistical Association. Earlier NCVS data had shown transgender respondents reporting violent victimisation at roughly 2.5 times the rate of cisgender respondents.

researchers computers data

The Lancet audit

In July 2025 a research letter in The Lancet put numbers on the quieter half of the problem. Janet Freilich of Boston University School of Law and Aaron Kesselheim of Harvard Medical School compared 232 federal datasets modified between 20 January and 25 March 2025 against archived versions. They found 114 substantially altered. In 106 of those, the word gender had been swapped for sex. Fifteen carried any note that a change had been made at all.

One case stands in for the pattern. A Department of Veterans Affairs dataset on veteran health-care use in 2021, untouched since it was published in 2022, was amended on 5 March 2025: a column headed gender became sex, and the same swap ran through the title and the description. As of 1 May the dataset’s change log was still empty. The Journalist’s Resource summary of the study notes that the alterations spanned HHS, the CDC and the VA, and that respondents answer questions about gender differently from questions about sex — which is precisely what makes the substitution consequential rather than cosmetic.

The audit’s central finding was procedural. Federal statistical agencies operate under documentation standards requiring version histories, changelogs and metadata records for any modification to a public dataset. In most of the flagged cases, those records were absent.

A silent modification is a different thing from a deletion. A deletion is legible. A silent change is not.

The database that is not coming back

The starkest single removal, from a public-safety standpoint, was the National Law Enforcement Accountability Database — the only centralised online repository of federal law-enforcement officer misconduct. By September 2024, 94 federal agencies had submitted 4,790 records covering 4,011 officers and dating back to 2018, and hiring officials had queried it close to 10,000 times in its first eight months to screen candidates for federal policing jobs.

Its origins were bipartisan: a 2020 executive order signed by President Trump and a 2022 executive order signed by President Biden. Trump revoked the Biden order on 20 January 2025, and the Justice Department decommissioned the database four days later.

Federal hiring personnel now have no equivalent centralised source for cross-agency misconduct records. The National Decertification Index still runs, but it covers state and local officers, not federal ones.

Thinner staff, and why an archive is not the same thing

Every incoming administration edits federal websites — climate pages were rewritten in 2017, terminology shifted back in 2021. What came down in 2025 was not mainly rhetorical framing on policy pages. It was underlying statistical products: survey instruments, microdata files, dashboards, and the metadata that lets researchers reproduce an analysis. Terminology on a landing page can be rewritten by the next administration. A discontinued survey wave cannot be reconstructed after the fact. If the 2025 round of the National Crime Victimization Survey does not ask about gender identity, that data point does not exist for 2025, whatever any future administration decides.

Building a longitudinal dataset takes decades. Ending it takes an email.

Staffing compounded the problem. Pierson estimated attrition of 15 to 40 per cent at some statistical agencies following the across-the-board workforce reductions initiated by the Department of Government Efficiency. Fewer statisticians means fewer field surveys.

Erica Groshen, former Bureau of Labor Statistics commissioner and now a senior economic adviser at Cornell’s ILR School, told Marketplace that BLS has concentrated what it has left on the headline figures — the monthly employment report, the top-line consumer price index — at the cost of the granular tables specialists use. She named the American Time Use Survey, the National Longitudinal Surveys, occupational injuries and illnesses, and the Job Openings and Labor Turnover Survey as products that could be scaled back or cancelled if funding pressure continues. Michael Strain of the American Enterprise Institute pointed out what rides on the CPI: it indexes Social Security payments, and a small mismeasurement compounds into hundreds of billions of dollars.

The rescued copies carry limitations that federal-held data does not. A researcher submitting a paper can cite a CDC URL and expect a reviewer to accept it; a citation to a mirrored copy on a university server invites questions about provenance, chain of custody, and whether the file has been altered. This is exactly why the Library Innovation Lab signed its archive — though a signature attests that a copy matches what was collected, not that the collection was complete.

What happened in practice looked nothing like a funded archival programme. It was graduate students working evenings, librarians running scripts on shared infrastructure, and a Slack channel deciding which URLs to prioritise before the next round of takedowns.

Some of it came back. The Youth Risk Behavior Surveillance System is online again, under a court order and a disclaimer the department attached itself. NOAA’s disaster database is running again, maintained by a nonprofit instead of a federal agency. The law-enforcement misconduct records are not running anywhere.

What no archive restores is the collection itself. A survey that stopped asking a question in early 2025 has a hole in that variable from that date forward, and no amount of scraping fills it. Researchers a decade out — studying adolescent mental health, or federal law-enforcement misconduct, or the demographics of the federal workforce — will hit a discontinuity and find it timestamped to the winter of 2025.

The copies made that Friday afternoon in Kresge G1 are still sitting on university servers, signed and catalogued, waiting for someone to need them. Nobody in that room knew which files would turn out to matter. That was the point of grabbing all of them.