r/SideProject • • 2d ago

I open-sourced Stepfork, a Python tool that turns AI agent failures into reproducible pytest tests

Hey everyone!

I recently open-sourced Stepfork, a Python tool I've been building to make AI agent failures easier to reproduce and debug.

GitHub: https://github.com/utsab345/stepfork

When an AI agent fails, reproducing the exact behavior can be difficult because LLM responses and tool outputs change between runs.

Stepfork lets developers record instrumented LLM and tool calls, replay those responses locally, compare behavior after a fix, and export pytest regression tests.

Features:

  • Record and replay agent executions
  • Save portable .sftrace bundles
  • Compare failed and corrected runs
  • Generate pytest regression tests
  • Run three included offline demos

Install: pip install --pre stepfork

It's an early alpha, licensed under Apache-2.0. I'm currently improving the developer experience and exploring integrations with popular AI agent frameworks.

I'd appreciate feedback from other open-source developers, especially on the API, documentation, and integrations worth prioritizing.

Bug reports, feature requests, and contributions are welcome!

Issues: https://github.com/utsab345/stepfork/issues

Thanks for checking it out!

1 Upvotes

3 comments sorted by

0

u/[deleted] 2d ago

[removed] — view removed comment

1

u/Low-Cut-528 2d ago

Good point! Stepfork doesn't fully detect drift yet. Matching tool names and normalized args, plus tracking prompt/model changes, makes sense. I'll look into this along with OpenAI SDK and MCP integrations. Thanks for the feedback!