TURNCOAT.agent-sandbox · Beehive AI Labs
HomeDashboardLive benchmark ↗
Hands-on assessment · one agent workflow

Find out what an injected instruction makes your agent do.

Before your next deployment, test whether untrusted messages, documents, or tool outputs can steer your agent into an unauthorized action. We run the assessment with you and deliver the evidence.

$500 USDone-time · tax included
1 workflowone pinned model and version
1 retestafter your fixes

Opens an email to arrange a free 20-minute call. We confirm fit, scope, and delivery date before payment. Already agreed? Pay for your assessment.

The pilot

A small test with a useful answer.

Included for $500

  • One agent workflow and one agreed test environment.
  • Eight agreed attack cases, each paired with a normal-task control.
  • Three trials per case and control: 48 scheduled runs.
  • A private report with observed tool calls, findings, reproduction steps, and prioritized fixes.
  • A findings walkthrough and one rerun of the same matrix after your fixes, within 30 days of the report.

Agreed before payment

  • The normal task, actions your agent is allowed to take, and what counts as a failure.
  • A runnable test endpoint or image and up to one hour of setup assistance.
  • Synthetic data and canary credentials; production access is not needed.
  • You supply model access with an agreed spending limit. Provider usage is separate from the $500 fee.
  • Custom integrations and broader testing are scoped separately. We confirm compatibility first.

How it works

1. Scope together

In 20 minutes, identify one meaningful workflow and its boundaries. We agree on the test cases, access, cost limit, and delivery date in writing.

2. Test and explain

We handle the run and walk through the report. Each finding separates the action requested by the attacker, the action issued by the agent, and any side effect we actually observed.

3. Fix and retest

Your team makes the changes. We rerun the agreed matrix and check the normal tasks too, so a blocked attack does not hide broken functionality.

Evidence you can inspect

Start with our public benchmark.

Our SWE-agent benchmark publishes the payload methodology, observed attempt rates, incomplete runs, and supporting evidence. It includes corrected findings and supplemental runs. Those results describe that tested configuration; your assessment measures your own workflow.

Read the benchmark and its limitations →

An issued tool call is not proof of a completed external action. We report side effects only when instrumented evidence supports them, and label incomplete runs. A passing test is evidence about the tested cases, not a guarantee of security or a certification.

Who this fits

Founders and engineering leads deploying agents that read external content and can execute code, update records, send messages, or access private data. Bring a concrete workflow and an upcoming release, model change, or customer security question.

You do not need to learn the dashboard first. We work directly with your technical contact. Customer findings stay private; any case study or reference requires separate permission.

One workflow · $500 USD · one retest

Bring the workflow you want to trust.

Tell us what the agent reads, what it can change, and when you plan to ship. We will confirm whether this pilot fits before you pay.

Prefer to write directly? Email hello@beehivewebstudio.com. If we already offered you a free design-partner assessment, that offer still stands.

Ready for payment

Scope agreed? Complete your booking.

Pay only after we have confirmed your workflow, compatibility, scope, and delivery date in writing. The assessment is $500 USD total, including applicable tax, with no subscription. Model-provider usage remains separate under your agreed spending limit.

Secure checkout hosted by Stripe. If we offered you a free design-partner assessment, you do not need to pay.