api p95 latency 2.4s, threshold 800ms. Webhook delivered, workflow started.
AI agents with scoped permissions.
An agent is a prompt, the tools it may call, the skills it can read, the shape of the answer it returns and the identity it runs as. Use the built-in agents or define your own.
Free forever plan, no commitment, in your own cloud account.
Built from the real interface. Open the agents view in the live demo. No sign-in.
- Today
- 51 minMedian time to resolve a high-impact outage. 39% take over an hour.New Relic, Observability Forecast, 2024
The median MTTR for high-business-impact outages was 51 minutes; 39% of respondents said MTTR was an hour or more
1,700+ technology professionals in 16 countries - With Wayfinder
Wayfinder timings are from our reference deployment and demo flows, not from customer studies.
- Under 15 minFrom the alert to a tested fix in a pull request. Timed on our reference deployment.
- WFWorkflow cost-control, run 88Input: CostFinding, api, +34% week on week. Identity: sa:cost-triage. Sandbox: own kernel.
- AICalled read_metrics(api, 7d)CPU 11% average across 6 replicas. Requests flat. Replica count changed on 14 Sep.
- AICalled list_resources(stack: payments)api: 6 replicas, requested by deploy #2231. Previous: 3.
- AIResultCause: replica bump left in place after load test. Proposed: 6 to 3, saves 34%. Tested in a replica, p95 unchanged.structured result2 tools calledheld for approval
Before and after
Before
- Agents running on laptops, one per engineer
- Prompts passed around in a repo
- Personal MCP setups nobody else can use
- No identity, no limits, no record
- Every incident still a person at 3am
With Wayfinder
- Agents versioned and published for every team
- Tools and MCP servers approved once, reused everywhere
- Identity scoped per workspace and environment
- Own kernel, default-deny network, no route in
- Incidents triaged with root cause before anyone is paged
How it works
Define it
Prompt, tools, the skills it may read, MCP servers, the result schema and the service account. As configuration, applied with wf apply.
Run it
Interactively from the UI or CLI, or hands-off from a workflow task.
Read the run
Every conversation records the tools called, the evidence used and the structured result.
What an agent looks like
The incident-triage agent that ships in the catalogue, trimmed. The comments are from the file.
# Incident triage: the shape of an agent you install rather than one that is
# simply there.
#
# The difference is the credential. This one reads a tenant's PagerDuty, and
# nobody can supply that token on the tenant's behalf. So it is offered in the
# catalogue, installed as the tenant's own copy, turned on, and given their
# credential at usage.
apiVersion: inference/v3
kind: AIAgent
metadata:
name: incident-triage
spec:
# It reads Wayfinder to triage an incident, so a tenant grants it viewer.
suggestedRoles:
- viewer
applicableScopes:
- tenant
agentDescription:
description: >-
Triages a PagerDuty incident against the Wayfinder resources behind it.
Read-only: reports what it found and changes nothing.
userInputsNeeded:
- the incident number or title
example: Triage incident 4821 and tell me which environment it affects.
# Answered at usage, not at install. A credential is only ever presented to
# the host it was created for.
inputs:
- name: pagerdutyCredential
type: externalidentity
identityTypes:
- BearerToken
Skills teach an agent how you work.
A skill is written knowledge an agent reads when it needs it, not another prompt to maintain. Write how your organisation names buckets, which regions are allowed, how a release goes out, and point the agents that should follow it at the skill.
- Markdown you write,with up to 16 files an agent can open one at a time, so a long standard stays readable.
- Tools of its own,so reading a skill gives the agent the calls that go with it.
- Named per agent,so an agent reads the skills you allow it and refuses the rest.
The built-in agents are skills of the Assistant, so the same mechanism that ships the platform's own knowledge carries yours.
In the product
- Deploy 14 min earlier bumped api to 6 replicas
- Connection pool saturated on postgres
- Fix proven in a replica: pool 20 to 60
Raise api connection pool to 60. Evidence attached. Summary posted to #payments-oncall and Jira PAY-881.
- Waiting for a human to merge
What this shows
- The alert fired a workflow; the workflow ran the agent with a scoped identity.
- The agent read metrics and resources through approved tools, then proved the fix in a replica.
- The write, a pull request, went through a gated workflow task, not the agent.
Components
The definition: prompt, tools, skills, result schema, identity.
Written knowledge an agent reads when it needs it: a body, up to 16 files and the tools that go with them.
One run and its full transcript.
Built-in observation tools, your MCP servers, the platform's.
For agents with a shell: a container with its own kernel, files kept between turns.
Wayfinder's own, or point at models you control.
Plan editors, the incident triage agent, the Wayfinder assistant.
Security and governance
An agent acts as the service account you name, whoever started it.
The tools you grant are the permissions; leave writes to a gated workflow task.
Agents with a shell get their own kernel on nodes that exist only while they run.
Every tool call is attributable and every conversation can be opened again.
Try an agent on one alert.
Start free, connect one webhook and read the first run.