Helix Hyperstrategies LLC

Let your agents run wild — inside a box you control.

We build agent sandboxes: hermetic, instrumented environments where autonomous agents can do their worst, so teams can ship AI systems they actually trust — not ones they just hope behave.

blast radius: 0replay: deterministic ✓

fig. 01 — agent under observation, fully contained

agent sandboxinghermetic environmentsdeterministic replayfault injectionred-team arenaspolicy guardrailsblast radius: zerobehavioral evalsagent sandboxinghermetic environmentsdeterministic replayfault injectionred-team arenaspolicy guardrailsblast radius: zerobehavioral evals

What we do

One problem. One obsession: no surprises in production.

THE AGENT SANDBOX

A padded room for your agents.

Autonomous agents are great at finding the one thing you didn’t want them to touch. Ours run in isolated, instrumented environments where they can do their worst — and you get a full recording of it.

Hermetic execution environments

Agents browse, code and call tools with zero blast radius. Nothing leaves the box unless you let it.

Deterministic replay & run diffing

Weird behavior becomes reproducible instead of folklore. Re-run the exact same episode, change one thing, see what moves.

Tool and API mocking with fault injection

See how your agents behave when the world breaks — timeouts, bad payloads, half-finished writes.

Policy guardrails and audit trails

Every action your agents take is scoped, logged and reviewable — the kind of record your security team will actually sign off on.

From the lab

Can an agent build its own harness? We measured it.

Everything we build rests on one claim: you can’t improve an agent you can’t re-run. So we tested it on ourselves. We put a small product model on 40 internal SRE triage tasks inside our own sandbox — fixed task set, deterministic scoring, every episode replayable — then changed exactly one variable: who designed the harness around the model. It designed one for itself; a frontier model got the identical tasks and the identical failure feedback and designed another. Same model doing the work underneath, both times. The harness moved the number by 25 points.

fig. 02 — can a model build its own harness?

held-out accuracy · 40 internal SRE triage tasks (20 train / 20 held-out) · product model qwen3:4b · deterministic scoring

no harnessdesigned by the model itselfdesigned by Claude
no harness (baseline)no harness (baseline) · 0.50 held-out accuracy0.50designed by the model itselfdesigned by the model itself · 0.60 held-out accuracy0.60self + training feedback ×3self + training feedback ×3 · 0.65 held-out accuracy0.65designed by Claude, same feedbackdesigned by Claude, same feedback · 0.85 held-out accuracy0.85

The frontier-designed harness beat the model’s self-designed one by 25 points on held-out tasks — with the same product model underneath.

See the numbers
conditiondesignerfeedbackheld-out
baselinenone0.5
self-designedqwen3:4bnone0.6
eval-iteratedqwen3:4btrain failures ×30.65
frontier-iteratedClaudesame feedback, same iterations0.85
fig. 03 — identical feedback, different designer

training-task accuracy per revision round · same failure feedback to both designers

qwen3:4b revises its own harnessClaude revises the same harness
0.500.751.00round 0round 1round 2round 3qwen3:4b revises · round 0 · 0.55qwen3:4b revises · round 1 · 0.65qwen3:4b revises · round 2 · 0.60qwen3:4b revises · round 3 · 0.600.60 — plateauClaude revises · round 0 · 0.55Claude revises · round 1 · 0.75Claude revises · round 2 · 0.950.95

Given the exact same failure feedback, the subject model plateaued while the frontier designer climbed monotonically. Feedback isn’t the moat — the intelligence consuming it is. Either way, the thing being tuned here is the harness, not the weights.

See the numbers
roundqwen3:4bClaude
00.550.55
10.650.75
20.60.95
30.6

a.

Same product model, two harnesses, 25 points apart on held-out tasks. What surrounds the agent matters more than which agent you picked.

b.

More feedback did not close the gap: three rounds of train-failure feedback moved the self-designed harness 0.60 → 0.65, while the same feedback carried the frontier-designed one to 0.85.

c.

None of this is visible without a sandbox. A fixed task set, deterministic scoring and replayable episodes are the only reason a 25-point swing can be pinned on the harness instead of on luck. That instrumentation is the product.

Fine print: this is an internal pilot instrument, not a paper. 40 internal SRE triage tasks split 20 train / 20 held-out, product model qwen3:4b, deterministic scoring, and the same failure feedback and revision budget given to both designers. Small task counts mean wide error bars — treat the direction as the finding, not the decimals. We re-run the instrument on each new model generation.

How an engagement runs

Boring process, exciting findings.

1

Scope

A deep-dive on your stack, your agents and what keeps you up at night. We leave with a threat model, you leave with a plan.

2

Instrument

We stand up the sandbox around your agents and wire in replay, mocking and audit trails. Your infra, our tooling.

3

Harden

Findings become guardrails, evals and dashboards your team owns. We hand over keys and documentation — not a dependency on us.

Contact

Let’s put your agents in a box.

Tell us about your agents and the systems they can reach. We’ll tell you what we’d poke at first — the first conversation is on us.

Helix Hyperstrategies LLC · Working with teams worldwide