Batch processing lets you run the Mibyan agent across hundreds or thousands of prompts in parallel, generating structured trajectory data. This is primarily used for training data generation — producing ShareGPT-format trajectories with tool usage statistics that can be used for fine-tuning or evaluation.
Overview
The batch runner (batch_runner.py) processes a JSONL dataset of prompts, running each through a full agent session with tool access. Each prompt gets its own isolated environment. The output is structured trajectory data with full conversation history, tool call statistics, and reasoning coverage metrics.
Quick Start
Predictable cost at scaleBatch runs spin up many concurrent agent sessions, each making model calls and tool calls. A Nous Portal subscription bundles model access plus web search, image gen, TTS, and cloud browsers under one bill — useful when you want stable cost-per-trajectory without juggling rate limits across five vendor accounts. Set up with mibyan setup --portal, then point --model at a Nous model.
The input dataset is a JSONL file (one JSON object per line). Each entry must have a prompt field:
Entries can optionally include:
image or docker_image: A container image to use for this prompt’s sandbox (works with Docker, Modal, and Singularity backends)
cwd: Working directory override for the task’s terminal session
Configuration Options
Provider Routing (OpenRouter)
Reasoning Control
Advanced Options
Each prompt gets a randomly sampled set of toolsets from a distribution. This ensures training data covers diverse tool combinations. Use --list_distributions to see all available distributions.
In the current implementation, distributions assign a probability to each individual toolset. The sampler flips each toolset independently, then guarantees that at least one toolset is enabled. This is different from a hand-authored table of prebuilt combinations.
All output goes to data/<run_name>/:
Each line in trajectories.jsonl is a JSON object:
The conversations field uses a ShareGPT-like format with from and value fields. Tool stats are normalized to include all possible tools with zero defaults, ensuring consistent schema across entries for HuggingFace datasets compatibility.
Checkpointing
The batch runner has robust checkpointing for fault tolerance:
- Checkpoint file: Saved after each batch completes, tracking which prompt indices are done
- Content-based resume: On
--resume, the runner scans existing batch files and matches completed prompts by their actual text content (not just indices), enabling recovery even if the dataset order changes
- Failed prompts: Only successfully completed prompts are marked as done — failed prompts will be retried on resume
- Batch merging: On completion, all batch files (including from previous runs) are merged into a single
trajectories.jsonl
How Resume Works
- Scan all
batch_*.jsonl files for completed prompts (by content matching)
- Filter the dataset to exclude already-completed prompts
- Re-batch the remaining prompts
- Process only the remaining prompts
- Merge all batch files (old + new) into final output
Quality Filtering
The batch runner applies automatic quality filtering:
- No-reasoning filter: Samples where zero assistant turns contain reasoning (no
<REASONING_SCRATCHPAD> or native thinking tokens) are discarded
- Corrupted entry filter: Entries with hallucinated tool names (not in the valid tool list) are filtered out during the final merge
- Reasoning statistics: Tracks percentage of turns with/without reasoning across the entire run
Statistics
After completion, the runner prints comprehensive statistics:
- Tool usage: Call counts, success/failure rates per tool
- Reasoning coverage: Percentage of assistant turns with reasoning
- Samples discarded: Count of samples filtered for lacking reasoning
- Duration: Total processing time
Statistics are also saved to statistics.json for programmatic analysis.
Use Cases
Training Data Generation
Generate diverse tool-use trajectories for fine-tuning:
Model Evaluation
Evaluate how well a model uses tools across standardized prompts:
Per-Prompt Container Images
For benchmarks requiring specific environments, each prompt can specify its own container image:
The batch runner verifies Docker images are accessible before running each prompt.