Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result.

But the comparison isn't straightforward.

OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: