ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration
IBM Research has introduced ScarfBench, a specialised benchmark designed to evaluate autonomous AI agents on enterprise Java framework migration tasks, published via Hugging Face. The benchmark tests the capability of AI models and agentic workflows to refactor, upgrade, and modernise legacy enterprise codebases across complex Java application frameworks. It establishes standardized criteria for measuring agent accuracy, code consistency, and autonomous migration performance in legacy enterprise environments.