A new open-source harness called BootLoops is enabling AI models to perform exact scientific calculations, bridging the gap between automated problem-solving and genuine research. Developed by Prof. Matthew Schwartz, a physicist at Harvard University and visiting researcher at Anthropic, the tool is designed to address "Claude-shaped problems"—tasks that leverage the specific strengths of current AI models while relying on human expertise for direction.
What Happened
Schwartz introduced BootLoops after realizing that using AI models like Claude as general-purpose researchers yielded mixed results. Instead, he focused on areas where AI excels at computation and pattern recognition. The source code for BootLoops is available on GitHub. According to Schwartz, the harness allowed Claude to find connections between disparate fields such as particle physics, ecology, population genetics, economics, and linguistics.
Over a three-month period, the team produced 36 manuscripts across 18 fields with 19 co-authors. The process began with particle physics, where Claude computed 30 integrals using BootLoops; fifteen reproduced known results, while fifteen were computed for the first time. The team then expanded the application to other disciplines. In ecology, the model solved a 20-year-old equation from neutral biodiversity theory. When applied to data from Barro Colorado Island in Panama, the results indicated that tree species composition is changing 4.5 times faster than the theory previously allowed. Ecologist James O'Dwyer subsequently helped refine these findings into a predictive model.
Other applications included analyzing 5.7 billion mutation pairs from the 1000 Genomes Project to find evidence for gene conversion, creating an AI data editor for economics journals that checked 4,452 replication packages, and building a word stress database covering 6,072 languages in collaboration with three linguists.
Why It Matters
The release of BootLoops highlights a shift in how AI tools are integrated into scientific workflows. Schwartz argues that the traditional model of applying for long-term grants for calculations is becoming obsolete, as AI can potentially solve such problems overnight. This rapid capability change raises questions about the future of PhD training and curriculum relevance, with Schwartz noting that courses like "Python for Engineers" are becoming less essential as AI models handle implementation tasks.
However, the harness does not replace the need for human oversight. Schwartz emphasizes that AI results often only become scientifically valuable when domain experts intervene to set the direction. He warns of significant limitations, including the model's tendency to declare victory prematurely, misjudge task duration, and brute-force calculations rather than finding elegant solutions. Automated checks are described as unreliable, with the possibility of correct calculations leading to wrong conclusions. Additionally, the projects were noted to be "compute- and token-intensive," and the model tends to gravitate toward established debates rather than genuinely new questions.
The Bottom Line
BootLoops demonstrates the potential of AI harnesses to accelerate scientific computation and cross-disciplinary analysis. While it enables the rapid generation of manuscripts and new computational results, Schwartz cautions that the scientific method remains dependent on human guidance, taste, and rigorous validation to avoid unrealistic expectations and ensure robust findings.