Let’s start with the number that makes the rest of the story hard to picture: 200 billion. That is roughly how many rows of data a machine-learning program built by a Pasadena high schooler had to work through. Each row is one moment a retired NASA telescope caught a flicker of infrared light somewhere in the sky. Stacked over more than a decade of scanning, they add up to a table nobody could read by hand.
The person who wrote code to comb through it is Matteo Paz. When the flickers were sorted, his program flagged about 1.5 million potential new objects. In 2025 that work won him the top prize in a national science competition.
We want to answer the obvious questions: what was the code actually doing, why “potential” is the word that matters, and how a teenager ended up doing it inside one of Caltech’s research centers.
What is 200 billion entries of raw infrared data?
The data came from NEOWISE, an infrared space telescope that scanned the entire sky over and over for more than ten years before it was retired. Every pass added more detections. Paz’s mentor, IPAC senior scientist Davy Kirkpatrick, put the scale plainly: the count was creeping up towards 200 billion rows in the table of every detection the survey had made.
Raw is the key word. This was not a tidy catalog of stars and galaxies. It was individual snapshots, the kind of firehose most projects trim down before they even start, because the full stream is so unwieldy. The value hiding in it is change over time. If you watch the same patch of sky again and again, some points of light stay steady and some brighten and dim. The ones that vary are often the interesting ones. Kirkpatrick’s original plan for the summer was modest: take a small piece of the sky and look for variable stars.
What was the code actually doing?
Paz’s program was built to catch tiny differences in infrared brightness across those repeated measurements. He described the method in a peer-reviewed paper in The Astronomical Journal, which he wrote alone. It combines signal-processing math with a neural network that learned to tell one kind of flicker from another.
The output was not just a pile of “this one changes.” The program sorted sources into a small set of categories, separating steady sources from the various ways an object’s light can vary. Some things pulse on their own. Some dim on a schedule because a companion star passes in front of them. Some flare once and fade. Sorting candidates into groups is what turns a raw list into something a scientist can actually search, because a researcher hunting for eclipsing pairs of stars does not want to sift through every distant flickering galaxy to find them.
What strikes us is the reach Paz claims for the approach beyond astronomy. He has said the model can be used for other time domain studies in astronomy, and potentially anything else that comes in a temporal format. That “potentially” is his, and worth keeping. Anything that arrives as a stream of measurements over time is a candidate in principle. Whether the model actually proves useful on, say, financial data or sensor readings is an open question, not a proven result.
1.5 million “potential” is not 1.5 million discoveries
The distinction the headline number can hide is this: the 1.5 million potential new objects are candidates, not confirmed finds. Some may turn out to be sources already cataloged. Some will be false alarms, quirks of the data rather than real varying objects. The catalog is a list of things worth a closer look, and the looking is a separate job.
This is not a knock on the work. It is how surveys of this kind function. A program built to trawl 200 billion measurements is valuable precisely because it narrows an impossible search down to a manageable set of leads. Among the candidates are objects whose brightness shifts over time, and the Society for Science’s summary of Paz’s census of infrared variable objects lists supermassive black holes, newborn stars and supernovae among them. Confirming any single one takes follow-up observation and analysis. Treating the 1.5 million as finished discoveries would overstate what the catalog is, and understate why a filtered list of leads is genuinely useful.
The paper describing the method was published in November, and Paz and Kirkpatrick have said they plan to release the full catalog so others can begin that follow-up work.
How did a high schooler end up doing this?
The short answer is a chain of outreach programs rather than a single lucky break. Paz’s route ran through Caltech public lectures and a summer research program that paired him with Kirkpatrick in 2023. The project was meant to last six weeks. Paz had something larger in mind from the start, telling Kirkpatrick on the first day that he was considering working toward a paper, a far bigger goal than six weeks would normally allow. By his account, the mentor did not discourage him.
What the mentorship bought, by Paz’s account, was less technical instruction than room to think. It left space for the ambitious version of the project to survive, and the six weeks became the start of a much longer piece of work.
What happens to a catalogue this size now?
The prize, first place and its $250,000 award in the 2025 Regeneron Science Talent Search, reads more like a starting line than a finish. Paz now works at IPAC as a Caltech employee, and the catalog he built is still a set of leads waiting to be checked.
Which of the 1.5 million candidates are real and new, which are already known, and which are just noise, are questions the release is meant to hand to the wider community. A candidate list is an invitation. What astronomers make of it, and what the method does when pointed at the next flood of survey data, is the work that has not happened yet.