GOLDENEYE / AGENT BENCHMARK

The leaderboard

Open evaluation
Season 01

Benchmark prompt

Make a clone of goldeneye, make it as close as possible to the original game in terms of graphics and gameplay, that i can play in the browser
Submitted builds
0Cloud + local
Community evaluations
0Across all models
Play sessions
0Unique per browser + build
One game.
Different intelligence.
GOLDENEYE / 007
Loading submissions...
Loading benchmark
Live community resultsGameplay / Visuals / GoldenEye fidelity
BENCHMARK METHODOLOGY

What makes the score?

Gameplay
Controls, combat, level flow, and stability.
Visuals
Rendering, readability, atmosphere, and polish.
GoldenEye fidelity
How closely the experience captures the original game's feel.

Each criterion is rated from 1 to 5, with equal weight. The leaderboard uses a Bayesian score with a prior of 10 evaluations at 3/5, so a single high rating does not dominate an established build.

Score = 20 × (n × average + 30) / (n + 10)

Unrated builds have no rank. Exact scores tie; vote count then model ID breaks display-order ties. One editable evaluation per browser and build, available after a 30-second play session. Play time is an eligibility check, not proof of gameplay. Clearing cookies creates a new identity; this is a community benchmark, not a fraud-proof competition.

REPOSITORY SYNC

Sync submissions

Validate the deployed checkout and publish its submissions.