Skip to content
The AI-First Web

In one AI-built sky map, a person caught what AI reviewers missed

Anthropic says Claude Science built a complete UV sky map over several days, a job researchers tend to put off. The checks matter as much as the speed.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
In one AI-built sky map, a person caught what AI reviewers missed
In brief
  • Anthropic researcher Brice Ménard reports that Claude Science built a full ultraviolet sky map, predicting the third of the sky no telescope had measured.
  • Two rounds of AI review missed a visible flaw. Ménard, a human, caught it. The map also labels each pixel as measured or predicted.
  • Leaders can ask the same of any AI output: what was measured, what was estimated, and who checks the checkers.

The work nobody had time for

Every organisation has a shelf of projects that would help everyone and belong to no one's top priority. One example comes from astronomy. Brice Ménard teaches astrophysics, holds a post at Johns Hopkins University and works as a researcher at Anthropic. He wanted a complete map of the sky in ultraviolet (UV) light for his classes. The only one he could show students was, in his words, full of holes.

The lesson here is that the first thing AI changes may not be the glamorous work. It may be the backlog: valuable tasks that were never worth a person's scarce weeks. Ménard writes that with Claude, such lower-priority work has become easier to tackle.

About one-third
Share of sky never observed in UV
Source: Anthropic, Brice Ménard research post, reported by The Decoder (October 9, 2026)

How the map was built

Earth's ozone layer blocks ultraviolet rays, so telescopes must orbit above the atmosphere to capture them. NASA's GALEX mission imaged about two-thirds of the sky between 2003 and 2013. It skipped places with very bright stars, including much of the Milky Way's plane, to protect its detectors.

Ménard says Claude Science directed a team of AI agents. Some searched the web for public UV surveys and downloaded them. Others made each survey consistent, including removing glare around bright stars. Another group cross-calibrated the surveys, put them at the same resolution and placed them on one coordinate system.

The hard step was the missing third. Claude used inpainting, a machine-learning method that learns how parts of an image relate to their surroundings and restores missing pieces. It also learned how UV brightness relates to visible, infrared and radio data. It applied that link where no UV data exists and estimated its own confidence at each point.

Individual stars came last. Using visible-light readings from the European Space Agency's Gaia satellite, Claude inferred UV output for more than 100 million of them. It placed those estimates over the smooth background.

Within about 10%
Accuracy on deliberately hidden regions
Source: Anthropic, Brice Ménard research post, reported by The Decoder (October 9, 2026)

The check that mattered

To test the method, Ménard had Claude hide parts of regions where UV data already exists, then fill them in blind. After several rounds of refinement, the estimates landed within about 10% of the real measurements. Ménard calls that difference almost imperceptible to the human eye.

The more instructive moment came later. One evening Ménard noticed faint circles in one of the dimmest fields, each slightly brighter or darker than its neighbours. They were the footprints of single GALEX observations, caused by leftover glow from Earth's atmosphere.

Claude had named this as a known issue before work began. Even so, two review passes by other agents signed off on the map, and the problem went uncaught. Ménard pointed out the circles. The agents traced the cause and corrected the glow across all 38,000 observations after a couple of hours of processing.

38,000
GALEX observations corrected after a human spotted the flaw
Source: Anthropic, Brice Ménard research post, reported by The Decoder (October 9, 2026)

What this means for an organisation

This is one account, written by an Anthropic researcher about Anthropic's product. The sources do not include independent review. It is a single case, not a trend.

Still, the case carries a lesson. In this case, the work was produced faster than reviewers, human or AI, caught its flaws. Agent reviewers passed an error that a person saw by looking at the output. Naming the issue early did not mean it was fixed.

The map also holds a good answer. Extra layers label every pixel as measured or predicted and give uncertainty estimates. A reader can see which parts are data and which are inference. Leaders should ask whether the AI-produced reports in their own organisation draw that same line.

Questions to put to your team

Which backlog items would be worth finishing if the cost fell sharply, and who owns the result once they are done? When AI produces a report, can a reader tell what was measured from what was estimated? Is there a hold-out test, like the hidden-region test here, that shows how wrong the estimates can be?

Finally, ask who looks at the finished output. Ménard says his own part was guiding the agents. In this case, that guidance included the catch that mattered. This case suggests human attention stays an important control when production gets faster.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Anthropic.

Share this insight