← All articlesENGINEERING

Benny, Cursor’s Overnight Bug-Fixing Bot

Lauren Tan built Benny to turn bug reports into reproduced issues, verified fixes, and reviewable pull requests while she was away from the keyboard.

The pain point Benny solves

Bug reports create two different jobs. Someone has to decide whether the report describes a real failure, and then someone has to fix it without breaking something else. Coding agents make the second job faster. They do not automatically make the first one trustworthy.

Lauren Tan built Benny around that verification gap. Cursor was receiving bug reports through its internal dogfooding channels, often with screenshots, videos, or partial descriptions. The slow part was reconstructing the failure, checking whether it was already fixed, and collecting enough evidence to make the next decision obvious.

Benny gives that queue a standing owner. The useful output is not a confident paragraph about what probably happened. It is a reproduced issue, a fix when the evidence supports one, and a review package that shows the result.

About Lauren Tan

Lauren Tan is an engineer at Cursor and a member of the React Core team. Her background includes engineering and engineering-management roles at Netflix and work on React Compiler at Meta. That mix matters here because Benny is designed like a managed engineering process, not a one-shot coding prompt.

In a recent interview with Roshan Sadanani, Tan described Benny as an internal Slack bot with a small dog avatar. The original goal was blunt: let agents work through Cursor bug reports while she slept. Questions about how Benny worked helped shape the thinking behind persistent Grok Bots with identities, computers, routines, and automations.

How Benny works

Benny starts with the report, then gathers the surrounding evidence before touching the code. Tan’s published workflow checks the codebase, asks for missing reproduction steps, and looks through relevant history and internal context to decide whether the behavior is a bug or an intentional product decision.

Once the failure is concrete, Benny uses a cloud computer to reproduce it. A consistent reproduction can move into a fix. Performance problems can trigger CPU traces or heap snapshots. The workflow then verifies the result against the ticket, records before-and-after evidence, and opens a pull request with that evidence attached.

  • Intake the bug report and collect the missing context.
  • Reproduce the failure before proposing a code change.
  • Fix only when the reproduction supports the diagnosis.
  • Run verification and preserve screenshots, video, traces, or logs.
  • Return a pull request for human review.

What is worth copying

The transferable part is the evidence contract. A bug-fixing bot should not earn trust because it writes code quickly. It earns trust when it can show the break, show the change, and show that the original failure no longer occurs.

Start with one narrow class of reports and one required proof package. Keep merging and production changes behind review. If the bot cannot reproduce the issue, it should return the missing information instead of manufacturing a diagnosis.

Sources