OY Labs posts perfect ARC-AGI-3 public score
OY Labs says its OY1-AGI harness achieved a 100% score on the public ARC-AGI-3 set using GPT-6 Astra, though the result is not a held-out certification of general intelligence.

A new benchmark result is highlighting how much AI performance can depend on the system wrapped around a frontier model, not only the model itself.
What happened
OY Labs announced that its OY1-AGI harness achieved a 100% score across the public ARC-AGI-3 benchmark using GPT-6 Astra at high reasoning. The company reported a run cost of $415.37 and has open-sourced its benchmark implementation. Its leaderboard application remains under review.
Why it matters
OY Labs is not claiming to have trained a stronger foundation model. Instead, it built a harness that guides an existing model’s problem-solving process. That illustrates how orchestration, memory and interaction design can materially change benchmark performance and cost.
The bigger picture
Public benchmark results need careful interpretation. OY Labs itself notes that its result does not establish general superiority on held-out tasks. As frontier models improve, evaluation will increasingly need to distinguish between model capability, system engineering and benchmark-specific optimisation.
