Skip to main content
BrowseComp is a benchmark for browser operation competition tasks, evaluating agents’ comprehensive browser operation capabilities.

Overview

Features

Competition-grade Tasks

Tasks from browser operation competitions with high difficulty

Comprehensive Skills

Tests a wide range of browser operation capabilities

Quick Start

Run Tasks

Evaluate Results

Data Loading

BrowseComp supports local JSONL files or HuggingFace downloads. To use HuggingFace:
The HuggingFace parquet file is converted to JSONL in the HF cache before use.

Evaluation Metrics

Data Format

Task data is stored in benchmarks/BrowseComp/data/: