VibeJam: A User Study Platform for Web Development with Agents
Abstract Overview
VibeJam is a browser-based platform designed for online user studies of agentic web development, where people collaborate with AI systems to build websites. Informed by a formative study with eight NLP and coding researchers, the system supports web-development tasks, agent customization, chat and plan modes, diff-based review, live website previews, and researcher-defined evaluation workflows. VibeJam uses the open-source Aider agent by default and is built to mirror contemporary agentic coding workflows without requiring participant setup. The paper positions the platform as experimental infrastructure for studying human-agent coding dynamics and releases both the platform and a suite of 55 game-based website creation tasks.
Novelty
The paper introduces an open-source, browser-based study platform specifically built for user-in-the-loop evaluation of agentic coding, moving beyond offline benchmarks and setup-heavy IDE plugins. It combines customizable agent configurations, plan and chat modes, interactive diff reviews, sandboxed live previews, and an integrated post-task annotation workflow for researchers.
Results
In a pilot study with five experienced AI programmers, participants found VibeJam fun, easy to use, and similar to commercial tools, reporting higher perceived quality and faster website construction than without AI. In an evaluation where 13 junior programmers built 22 game websites, user-plus-agent teams outperformed agent-only baselines on enjoyment, creativity, and style, while task adherence remained high across conditions.
Key Points
- VibeJam provides a zero-setup, browser-based experimental platform for evaluating human-AI collaboration in web development, complete with comprehensive interaction logging and built-in submission annotation.
- The platform mirrors modern coding tools through direct agentic code editing, dedicated chat and planning modes, git diff review, and live sandboxed website previews.
- Empirical studies showed that users found the system comparable to commercial agentic tools, and junior programmers collaborating with the agent produced higher-quality games than agent-only baselines across multiple subjective criteria.