Clear Runway
Tells you why your agent will get stuck before you hand it real work.
Last assessed · Methodology v5.1 · How ideas are researched and assessed
Problem in brief
Teams building LLM agents that act inside CRMs, helpdesks, finance tools and email find that agents which demo well break in production for environmental reasons: the service account is missing a scope, an endpoint behaves differently than documented, or a company name in one system does not match the same company in another. Public issue trackers show these failures surfacing opaquely — an insufficient_scope refusal reported to the user as an expired token or an unreachable server. Because nobody knows in advance where it will break, teams babysit each step and try to patch it with prompt tweaks, and the handed-over task comes back as a stream of approval requests.
- Who has this problem
- Engineering teams building LLM agents that act inside company systems like CRMs, helpdesks, finance tools and email
- How they cope today
- Supervising each step manually or trying to fix behaviour through better prompt engineering
- What it costs them
- Friction and annoyance for users, and expensive agent deployments that underperform expectations
- Opportunity explored
- Clear Runway, described below

The problem in detail
Agents that demo well fail in production for boring environmental reasons: the service account lacks a scope, an integration endpoint behaves differently than documented, or the agent can't match 'Acme Corp.' in the CRM to 'ACME Corporation' in the invoice. Because nobody knows in advance where it will break, teams babysit every step and try to patch it with prompt tweaks, so a task that was supposed to be handed over comes back as a stream of approval requests.
Proposed product
A pre-flight checker that sits between the agent and its tools. You point it at your agent's tool definitions and the credentials it will actually run with, and it dry-runs every capability against the live (or sandbox) systems: checks each OAuth scope and object-level permission, calls each endpoint with representative arguments, and runs the agent's entity references against real records to flag names, IDs and emails that resolve to zero or multiple matches. Output is a blocker report ordered by likelihood of derailing the task, with the exact fix (grant this scope, add this alias, this field is read-only for this role). At runtime it holds the same checks as a gate, collecting anything genuinely ambiguous into one batched review at the start rather than interrupting the user step by step.
Why now
Agent clients are now failing in ways that are documented in public issue trackers: GitHub issues on Claude Code and fastmcp show tool calls dying on HTTP 403 'insufficient_scope' and surfacing as 'token expired' or an unreachable server, so the real cause is hidden from the team. At the same time review synthesis of Salesforce Agentforce (4.3/5 across 1,205 G2 reviews) reports that data quality, permissions and prompt configuration are all left to the customer — the environmental setup is the recurring complaint, not model quality.
Smallest useful version
A CLI plus small Python library covering three connectors (Salesforce or HubSpot, Google Workspace, Slack). Ingest an agent's tool schema, take the service-account token, and emit a permission and entity-resolution blocker report before the agent is given real tasks. No runtime gating yet — the report alone should let a team predict and prevent most failed runs.
Potential ways to charge
- Open-source CLI — £0: scope and permission checks for the three connectors, run locally in CI (free tier to earn the paid product, standard for developer tooling)
- Team — £250 per month: hosted runs, endpoint dry-runs, entity-resolution checks against live records, history and diffing between releases (typical for comparable developer tools; evidence has no pricing)
- Platform — £1,200 per month: multiple environments and connectors, runtime gating with one batched ambiguity review, role-based service-account matrices (typical for comparable tools)
- Pre-launch audit — £4,000 one-off: hands-on run against a customer's agent and systems, delivered as a ranked blocker report with fixes (anchored to consultancy-style engagement pricing; evidence has none)
Routes to early customers
- Comment on and follow the open GitHub issues already describing insufficient_scope failures (Claude Code #92518, #44652, fastmcp #5237, coco-mcp #72) and contact the reporters directly
- Salesforce and HubSpot implementation partners and consultancies already publishing agent permission-troubleshooting content, who hit these blockers on every deployment
- MCP server authors and maintainers, who need a way to show their server works under real service-account scopes
- Agentic-AI freelancers and contract engineers advertising OAuth/API integration and guardrail work on freelance marketplaces — they carry the pain into each client
Main risks
- The core assumption is untested: if most production task failures come from model reasoning rather than statically detectable permission, endpoint or entity issues, the report becomes another dashboard nobody acts on.
- Platform vendors are absorbing the problem — Salesforce already ships agent-user access configuration and a beta 'Agent for Setup' permission troubleshooting tool, which could make per-platform scope checking free.
- Adjacent vendors already ship pre-execution credential validation (connection-status APIs) and MCP testing guides already treat expired tokens and missing scopes as first-class tests, so differentiation must rest entirely on the entity-resolution and endpoint dry-run layers.
- No evidence of a budgeted job to be done: no freelance posting was found paying anyone to hand-verify agent scopes or reconcile entity references, so buying may be absorbed into general integration work.
- Running dry-runs with live production credentials against real CRMs and mailboxes raises write-safety and security objections that could stall procurement.
Questions to test first
- Take five teams' agent logs from failed production runs and classify each failure: statically detectable environment issue versus model reasoning versus rate limits. This directly tests the riskiest assumption before any building.
- Reach out to the reporters of the four named GitHub issues and ask what they did to diagnose and fix the scope failure, and how long it took.
- Ship the free CLI with Salesforce scope and object-permission checking only, and measure how many blockers it finds per repo and whether teams fix them.
- Run a paid one-off pre-launch audit for two teams by hand, to see whether the findings are worth money and which check class produces the most value.
- Compare output against what Salesforce's 'Agent for Setup' beta and existing MCP scanners already report, to see how much unique ground the entity-resolution layer covers.
Evidence and sources
The sources this assessment is based on. Evidence level: Supported, counted from these sources as described in how ideas are assessed. Findings are what a source shows; anything the research only inferred is marked as an inference.
Research findings
SupportsSweep blog post '5 Salesforce Errors That Break Agentforce'; Salesforce Help docs on agent user access and 'Agent for Setup' permission troubleshooting · 2 December 2025
Shows: Vendor/consultancy troubleshooting content on hidden permission requirements breaking production agents
A vendor post states permission requirements for agent actions are often hidden (e.g. an action querying Contacts requires the Agent User to have Read on Contacts) and produces cryptic errors, noting an AI agent cannot work around a failure the way a human user can. Salesforce itself ships agent-user access configuration pages and a beta 'Agent for Setup' permission-troubleshooting tool — evidence the object-level-permission check the candidate proposes is a recognised pain, but also that platform vendors are beginning to absorb part of it (a differentiation risk for single-platform coverage).
SupportsGitHub issue trackers (anthropics/claude-code, PrefectHQ/fastmcp, mcp-zap-server, coco-mcp)
Shows: Issue-tracker reports of the exact failure mode (missing-scope tool calls) in widely used agent clients/servers
Multiple open GitHub issues show agent tool calls failing on HTTP 403 with WWW-Authenticate 'insufficient_scope': Claude Code reports it to the user as 'token expired' when the token is not expired (issue #92518), a separate issue reports step-up re-authorization never triggering (#44652), fastmcp issue #5237 states the client SDK only re-authorizes on a 403 insufficient_scope so 'a client never re-authorizes for the missing scope; it just sees a failed tool call', and coco-mcp #72 reports an insufficient-scope refusal surfaced as an unreachable server. This is observed, non-community-forum evidence that scope/permission gaps break agents opaquely at runtime — the core premise of the pre-flight checker.
Shows: Verified-user review synthesis of a major agent platform citing permissions/data quality as the recurring complaint
A review analysis of Salesforce Agentforce (cited as 4.3/5 across 1,205 G2 reviews) reports the recurring user complaint is not that agents fail per se but that data quality, prompt configuration, permissions and consumption forecasting are all left to the customer. This corroborates the candidate's framing that production agent failure is environmental/configuration-driven rather than model-driven, from a product-review source rather than a community thread.
ContextScalekit blog, 'Tool Call Failures in Production: Debug and Prevent Guide'
Shows: Industry guide enumerating tool-call failure taxonomy and pre-execution credential checking
An auth-infrastructure vendor guide describes ungranted scopes as requiring a new consent flow rather than a retry, describes a connection-status API used to check auth state before executing critical operations, and cites 60% of observed production LLM call errors coming from rate limits. Supports the problem but also shows adjacent vendors already shipping pre-execution credential validation, meaning Clear Runway's differentiation must rest on the entity-resolution and endpoint-behaviour dry-run layers rather than on scope checking alone.
Shows: Existing practitioner and vendor tooling in the 'pre-flight / gate the tools' space
Signadot's MCP testing guide treats auth as a first-class test concern — expired tokens, missing scopes and per-user permission boundaries get explicit tests run against real upstreams in an isolated sandbox per change. Separate posts describe dry-run modes and pre-flight gates for write-capable tools, and an issue thread benchmarks static preflight validation versus runtime authorization tests using scanners like mcpscan. The concept space is populated but fragmented (testing frameworks, security scanners, auth vendors); no result showed a single product combining scope checks, endpoint dry-runs and entity-reference resolution with a ranked blocker report.
Where the problem was reported
Public posts in which people described this problem, grouped into one problem before research began.
- Found many AI are expensive autocomplete instead of automation · r/projectmanagement · 17 September 2026
- Any positive experiences with Indeed Smart Assistant? (Sourcing) · r/recruiting · 15 September 2026
- Ask HN: Those running agents 24/7, what's your workflow and what are they doing? · news.ycombinator.com · 11 September 2026
- Why doesn't a Cursor for Word exist? · news.ycombinator.com · 10 September 2026
- A $10 Million Bet That Enterprise AI Agents Are Broken · dev.to · 30 August 2026
- Is "An agent with tools" the only valid LLM application? · news.ycombinator.com · 28 August 2026
Looked for, not found
- Freelance marketplace results showed general demand for agentic AI engineers with API/OAuth integration, guardrails, error handling and evaluation-harness skills, and one listing describing the need to make agent integrations reliable across CRM/ERP/external APIs with OAuth and webhooks. However, no posting was found that explicitly pays someone to hand-verify agent scopes/permissions or reconcile entity references pre-launch, and no keyword search-volume data was retrieved. Whether there is a distinct, budgeted 'pre-flight verification' job to be done — versus it being absorbed into general integration work — is therefore unknown.
This is a researched opportunity, not a guarantee of commercial success. The evidence level and sources show how much is known; the risks and questions to test show what is not.
Related product opportunities
Chosen for a shared category, customer or problem.
Developer tools
Dropbox for AgentsDevelopers increasingly run more than one AI coding agent across the same projects and machines.
Developer tools
Launch LaneSolo developers, indie hackers and open source maintainers can ship working software, but once it's live nothing happens.
Developer tools
Explain-Before-Merge: a comprehension gate for AI-written diffsDevelopers who ship most of their code through agents like Claude Code, Cursor or Copilot review diffs for whether they 'look right' rather than for whether they understand them.
Developer tools
Flight Deck: a worktree-per-agent control panel for Claude CodeDevelopers who try to run several AI coding agent sessions at once find that nothing manages the mechanics.
