Skip to content
← All case studies

PerpFlow cuts its agentic research workflow from 14 days to 12 hours with Sage

Claude Code delegates thousands of market classifications to Sage, keeping bulk results in files while Python tests the findings against predefined criteria.

A monumental marble researcher directs a sweeping stream of market cards through a towering arch, with blue ASCII textures carved across both forms.

Customer-reported workflow timing and research results

28×Faster research workflow

14 days → 12 hours · founder-reported elapsed time

5,972Distinct market cards

Classified against a fixed research taxonomy

85Pre-registered tests

Python evaluated the Phase 1 classifications

PerpFlow scans roughly 560 markets every four hours. It assigns each market a long, short, or neutral bias, scores the setup, and sends selected signals to users as Telegram cards. With Sage handling structured classification, founder Mo reports completing an agentic research workflow in 12 hours instead of 14 days.

A monumental marble researcher directs a sweeping stream of market cards through a towering arch, with blue ASCII textures carved across both forms.

The team connected Claude Code to Sage through a local Model Context Protocol (MCP) adapter to investigate missed LONG opportunities. Sage classified anonymized market descriptions. Python tested whether those classifications could improve signal selection under criteria fixed before the model runs.

I had to research about 6.3k structured market cases and run about 85 research tests. Without Sage, Claude would either have had to consume more tokens, spend more time reasoning through each individual market state, or substantially reduce the scope of the qualitative analysis.

Mo (@mochains)
Founder & CEO, PerpFlow

PerpFlow is in alpha, with signals delivered through Telegram and a private dashboard. The two views below come from its V3 update on X.

PerpFlow V3 equity-curve screenshot showing hourly mark-to-market performance from August 26 to September 25, 2026.
PerpFlow’s V3 equity curve, August 26–September 25, 2026. Historical product performance, separate from the Sage research results.
PerpFlow V3 open-book screenshot listing 22 positions, split evenly between longs and shorts, with profit-and-loss bars for each position.
PerpFlow’s V3 open book: 22 positions across 11 longs and 11 shorts. Select either screenshot to view it at full size.

Catch more opportunities without adding noise

In an audit of August and September 2026, PerpFlow’s production tiers captured 1,120 of 4,166 qualifying 24-hour LONG episodes, or 26.9%. A qualifying episode reached a volatility-adjusted upside target before an adverse move half that size.

The product problem was selective coverage. Expanding the watchlist could catch more upside moves while also sending users more failed setups. Any new rule had to improve coverage while satisfying predefined limits on precision, downside risk, and watchlist size.

Claude Code delegates classification to Sage

Claude Code prepared the research protocols, built the market cards, invoked Sage, and wrote the Python analysis. A local MCP server exposed health, single-decision, batch-decision, and JSONL classification tools. The bulk tool let Claude submit a file and a fixed schema.

Claude Code builds blinded cards. A local MCP adapter sends them to Sage for Tags, Choice, and Scale decisions, then saves the answers to JSONL. Python joins those labels to separately held outcomes and evaluates the frozen tests. Only counts, hashes, and retry information return to Claude.
Illustrative architecture adapted from PerpFlow’s technical report.

In Phase 1, each card contained information available at the scan time. Symbols, dates, outcomes, and sample groups were hidden. The key joining case IDs to market outcomes stayed outside the adapter’s allowed folder, so Python could evaluate labels after classification.

Sage decisionQuestion applied to each cardUse in the analysis
TagsWhich of 19 market archetypes apply?Test associations between defined market patterns and subsequent outcomes.
ChoiceWhich of 10 market families is primary?Compare fixed categories such as breakout preparation and unsafe countertrend bounce.
ScaleHow strong is this LONG-watch setup on a 0 to 4 rubric?Evaluate a consistent score alongside the categorical judgments.

Illustrative Python pseudocode for the local MCP tool. Parameter names are simplified; this is not PerpFlow’s production source or a Sage SDK example.

Example / python
# The full taxonomy is frozen before classification.
schema = read_json("research/frozen_schema.json")

summary = await mcp.call_tool(
    "sage_classify_jsonl",
    {
        "input_path": "research/blinded_cards.jsonl",
        "output_path": "research/sage_labels.jsonl",
        "schema": schema,
    },
)

# The adapter writes bulk answers to the output file.
# Claude receives counts, hashes and retry information.
agent_context.add(summary)

Bulk classification keeps the agent context compact

The adapter validated response shapes and wrote outputs to JSONL, recording the model name and schema hash. It returned counts, hashes, and retry information to Claude instead of placing thousands of classifications in the agent’s context. It refused to append results under a different schema.

Phase 1 wrote 6,288 rows: 5,972 distinct cards, 300 stability reruns, and 16 duplicate rows from an overlapping resume. The duplicates were outside the validation snapshot. That is the approximately 6.3k volume described by Mo.

Null answers remained undecided. In Phase 1, 22.7% of tag decisions were null. Python retained that category separately from negative answers. This preserved uncertainty for analysis; the report does not establish that those nulls reliably identified incorrect judgments.

Illustrative Python pseudocode after response decoding. Helper names and row fields are simplified; the statistical tests remain defined in the frozen protocol.

Example / python
labels = read_jsonl("research/validation_snapshot.jsonl")

for row in labels:
    row["tags"] = {
        tag: "undecided" if answer is None else answer
        for tag, answer in row["tags"].items()
    }

# Outcomes are joined only after Sage has classified.
cases = join_by_case_id(labels, sealed_outcomes)
verdict = evaluate_frozen_tests(cases, protocol)
archive_research_verdict(verdict)

The outcome for PerpFlow

The founder-reported business outcome is a shorter research cycle: an agentic workflow that normally took 14 days completed in 12 hours after Sage was added, a 28× speedup. The comparison measures elapsed workflow time. It is not a controlled model-speed benchmark or a measured reduction in token usage.

Sage handled repeated classification of market descriptions against a fixed taxonomy. It returned structured labels and scores for analysis; it made no trade execution decisions. Claude Code coordinated the workflow, and Python evaluated whether the classifications supported a change to signal selection.

The research did not establish a robust improvement to LONG selection, so no Sage-derived rule was deployed. The team reached a research conclusion sooner while retaining its validation requirements.

PerpFlow shows how a coding agent can delegate structured classification to Sage, retain results in files, and use Python to evaluate them against predefined criteria. To build a structured classification tool for your coding agent, see the Sage integration guide.

  • Coding agents
  • Structured classification
  • Research workflows

Build with Sage

Start with one decision.