Alexander Yue

Evaluations

Posts

How we built scalable evaluation infrastructure for AI web agents
Feb 23, 2026·Engineering

What LLM model should I use for Browser Use? The Definitive Browser AI Benchmark
Feb 19, 2026·Engineering

Browser Agent Benchmark: Comparing LLM Models for Web Automation
Jan 31, 2026·Engineering