Local runtime evidence for coding agents investigating performance, memory, execution, concurrency, and reliability.
Flameox coordinates maintained profilers, benchmark tools, and trace processors; preserves their native artifacts and provenance; and exposes bounded evidence to an agent. The agent forms the hypothesis. Flameox makes the measurements and experimental record inspectable.
It is not a profiler, hosted observability service, arbitrary shell or SQL gateway, source-code editor, or automatic bug finder.
Connect a supported MCP client through the guided setup:
npx flameox@latest setupRestart the client, open the project you intend to inspect, and ask it to:
Initialize Flameox in this project and list the available profiling capabilities.
Setup installs a versioned local runtime and changes only approved client
configuration. Project initialization is separate and creates .diagnostics/
only after the client calls the initialization workflow for its fixed project
root.
For source development:
uv sync --extra dev
uv run flameox init .
uv run flameox statusPython 3.12 or newer and the committed uv.lock are required.
symptom → capture or import → bounded evidence → hypothesis
→ discriminating experiment → supported, refuted, or inconclusive finding
Typical evidence sources include pyperf, py-spy, pytest-reportlog, coverage.py, Memray, Perfetto, torch.profiler, Nsight Systems, Nsight Compute, ROCprofiler, Compute Sanitizer, NVBench, and typed inference-provider exports. Availability depends on the host, permissions, installed extras, and selected adapter. Flameox reports missing evidence instead of silently substituting a weaker source.
A profile is exploratory. A performance or correctness conclusion requires a representative workload, declared metric and estimand, compatible run identity, preserved samples, and an appropriate semantic oracle.
Commands live in flameox.toml as argument arrays. Parameters are declared
scalars; there is no shell expansion.
schema_version = 1
[workloads.scan]
argv = ["python", "bench.py", "--implementation", "{implementation}"]
cwd = "."
timeout_seconds = 60
[workloads.scan.parameters]
implementation = ["baseline", "candidate"]
[workloads.scan.oracle]
strength = "cross_treatment_equivalence"
argv = ["python", "validate.py", "--implementation", "{implementation}"]
[experiments.scan_comparison]
workload = "scan"
design = "randomized_complete_blocks"
blocks = 10
treatment_factor = "implementation"
combination_policy = "cartesian"
primary_metric = "pyperf.workload"
polarity = "lower_is_better"
estimand = "median_paired_log_ratio"
practical_threshold = 0.05
confidence_level = 0.95
random_seed = 1984
[experiments.scan_comparison.factors]
implementation = ["baseline", "candidate"]The MCP configure_workload tool validates and writes the canonical definition
without executing it. A manually authored valid definition is active
immediately; there is no approval copy or secondary workload registry.
uv run flameox workload show scan --json
uv run flameox capture plan pyperf --workload scan \
--parameters '{"implementation":"baseline"}' --json
uv run flameox capture run pyperf --workload scan \
--parameters '{"implementation":"baseline"}' --jsonPlanning resolves every executable once. The resulting binding contains the
exact invocation path, canonical target, trust decision, and file identity.
Execution revalidates that binding instead of searching PATH again. Plans are
short-lived, single-use capabilities whose complete intent is retained in the
workspace SQLite control plane.
uv run flameox investigations create \
'{"question":"Does the candidate remove reverse-scan overhead?"}' --json
uv run flameox hypotheses record @hypothesis.json --json
uv run flameox experiment plan scan_comparison \
--investigation <investigation-id> --adapter pyperf --json
uv run flameox experiment run scan_comparison \
--investigation <investigation-id> --adapter pyperf --jsonExperiments retain randomized treatment order, attempted trials, failures, cancellations, validation receipts, and exclusions. Analyses resolve all input through one pinned corpus snapshot:
uv run flameox analyze hotspots <run-or-artifact>
uv run flameox analyze scaling <experiment-id>
uv run flameox analyze compare @comparison-request.json
uv run flameox analyze memory <run-or-artifact>
uv run flameox analyze execution <run-or-artifact>
uv run flameox analyze pytorch <run-or-artifact>
uv run flameox analyze failuresRead-only analysis does not create a durable claim. Use analyze record,
analyze record-comparison, or findings record when the result should become
part of the investigation history.
.diagnostics/ contains:
control-plane.sqlite3for plans, operations, runs, revisions, idempotency, and relationships;- content-addressed native artifacts;
- immutable Parquet generations and corpus commits;
- a rebuildable
catalog.duckdbanalytical cache.
Large evidence does not live in SQLite. Deleting catalog.duckdb does not
delete evidence; flameox catalog rebuild recreates it from committed
generations.
The CLI and MCP server expose bounded task-shaped operations, not shell strings, raw SQL, or arbitrary artifact bytes. Workloads may access the network unless active containment denies it. The control process performs network I/O only for explicit setup, upgrade, approved provider acquisition, or explicitly enabled symbol services—not during ordinary capture or analysis.
The trusted-local capture path records that descendant containment is not enforced. Projects that require managed containment can select it explicitly; planning refuses when the requested guarantee is unavailable.
uv run flameox --help
uv run flameox mcp serve --project-root .
uv run flameox mcp inspect --project-root . --jsonmcp inspect is the authoritative inventory of tool schemas, annotations, and
resource templates for the installed version. See CLI and MCP
boundaries for workflow and trust semantics.
uv run flameox validate
uv run flameox validate --full
uv run flameox catalog validate
uv run flameox catalog rebuild
uv run flameox recover
uv run flameox gc
uv run flameox gc --applyValidation never repairs evidence. Garbage collection is a dry run unless
--apply is supplied, and applied candidates first move to recoverable trash.
Permanent purge requires a separate explicit command naming an expired trash
manifest.
- Architecture — authoritative module and process boundaries
- Storage and evidence — authority, snapshots, and publication
- Investigations — experiments, analysis, and claim quality
- Adapters — producer ownership and compatibility
- Runtime safety — execution, filesystem, cancellation, and retention
- CLI and MCP — public workflow and trust boundaries
- Testing — suite ownership and CI lanes
- Contributing — development and pull-request workflow
uv sync --extra dev
uv run python tools/test.py list
uv run python tools/test.py core
uv run ruff check src tests tools
uv run mypy src tests tools
uv run pytest -qSee the testing guide before changing suite ownership, provider lanes, or collection topology. Flameox is available under the MIT License.
