import { Agent } from '@parcha/agentrun';
const agent = await Agent.create({
apiKey,
prompt: `Research companies, find their websites,
and return a brief with sources.`,
});
await agent.waitUntilReady();Self-improving agents that help you scale.
Grep agents learn from every run, turning repeated work into reusable code. Run them with your choice of models and scale your workflows while using fewer tokens.
# Set up Grep with my coding agent Help me turn repetitive work into reusable Grep agents, run them on new inputs, and improve repeated steps with workflows and code. Follow this guide in my current project. If an example task is included below, preserve it as my first task. ## Step 1: Read the docs and install the skills Start with these current sources: - Developer docs: https://grep.ai/developers - Coding-agent skill library (workflows for this host) and its installation: https://github.com/Parcha-ai/grepai-skills - Platform skill catalog (the skills a Grep agent runs with; a different thing from the library above): https://grep.ai/developers/skills - MCP connection and tools: https://grep.ai/developers/mcp - SDK and REST quickstart: https://grep.ai/developers/api/quickstart - Authentication: https://grep.ai/developers/api/auth - Live API schema: https://api.grep.ai/api/v2/openapi.json When two sources disagree, use this order: the live schema and a real response for API paths, fields, and status values; the repository README for the library's package name, install command, and skill list; the developer docs for everything else. This guide describes the procedure and yields to those sources on facts. Check whether the Grep skill library is already installed and which coding-agent host I am using. Read the repository README before installation. Node.js 18+ is required. If Node.js is missing, set up a supported Node.js LTS using this environment's normal version manager. Install only a published release. Take the npm package name from the repository's `package.json` (`name` field), and compare its `version` with `npm view <that name> version`. If the two differ, name both versions in your report. The README may also describe installing the repository checkout or an unreleased branch; do not use that path unless I ask for it. Then install or refresh the published release with the README's command for it. If the registry also has a package under another name, follow the README and mention the other name in your report. At the time of writing that command was: ```sh npx grep-research-skills@latest ``` Preserve existing user-owned skills and configuration. List the installed skills, compare them with the README, and check that this host discovers them; reload skills if needed. Read the installed SKILL.md files before using their workflows. The platform workflows cover routing, connection, turning repeated work into an agent, the agent lifecycle, and improving repeated work with workflows and code (at the time of writing: `grep-platform`, `grep-mcp`, `grep-agentify`, `grep-agents`, and `grep-optimize`). Take the names from the README and the installed list, not from this guide. If a workflow is missing from the published release, say so and continue with the documented MCP and API capabilities. You may read the repository's SKILL.md files without installing them; the scripts they reference are not on this machine in that case, so use the MCP tools or REST calls they map to. Do not invent a skill. ## Step 2: Understand my project and task If I have not already specified the context, ask whether I want to start a fresh project or integrate Grep into an existing one. For a fresh project, ask my preferred language and the repetitive task I want to automate. For an existing project, ask for its location if the current directory is not that project, then inspect its language, framework, package manager, directory conventions, environment loading, and existing agent integrations. Reuse these conventions. Use a supplied example task as the starting point. Otherwise identify a concrete repeated procedure from the work visible in this conversation or project, such as recurring commits, pull requests, or reports, and suggest making it reusable. Prefer a procedure whose inputs are text and whose sources are public; the agent cannot read this project or private systems unless you supply them as context. Ask only for missing inputs, required output, relevant source documents, and success criteria. Do not replace my task with a generic demo. ## Step 3: Connect Grep Read the current MCP setup instructions and configure this host's supported remote HTTP connector with: ```text https://api.grep.ai/api/v2/mcp ``` Preserve other MCP entries and any intentional custom Grep endpoint. Configure the connector for my user unless I ask to share it with the project. Complete the host's supported OAuth sign-in flow. If interactive authentication is required, give me the exact action to complete and continue preparing the integration while waiting. If this host cannot use MCP, or this session cannot complete a browser sign-in, say so and follow the authentication docs listed in Step 1: load a Grep API key through the project's secret manager or an ignored environment file, and ensure the running process actually loads it. Use the environment variable name that the consumer documents: the SDK docs name the variable the SDK reads, and the installed skill's scripts name the one they read. If two consumers need the key, set both names from one stored secret rather than inventing a third. Never ask me to paste credentials into this conversation or put secrets in tracked configuration. If keys exist for more than one deployment, ask which to use before creating anything. An agent belongs to the deployment and account that created it, and a connector to one deployment does not list agents from another. Record the deployment's API base with the agent reference. Verify the real connection: initialize the MCP client, discover its tools, and list agents I can access. For REST, use the equivalent documented read operation. Use live schemas for endpoint paths, argument names, and status values; do not guess endpoints or tool availability. Paths quoted in this guide or in older docs are examples to confirm against the schema. ## Step 4: Create or reuse an agent Use the installed agentify and agent-lifecycle workflows when available. First look for an existing agent that matches the procedure. Reuse it when appropriate; otherwise build an agent from my instructions, inputs, output requirements, and relevant context. Grant only the tools and skills it needs. If the live schema offers both a direct create endpoint and an asynchronous builder job, prefer direct create for a procedure you can already describe; it is immediate, and the builder is billed. The installed skill and the SDK quickstart may show only the builder; the live schema decides what exists, and this preference decides which to use. Run the documented validate or dry-run call first. It reports invalid or restricted skill and tool names before anything is saved. Read the stored configuration back and confirm it matches what you sent. Poll asynchronous builds using the returned ID. Inspect the completed configuration and lifecycle state; follow the documented activation flow if it is a draft. Do not create duplicate agents when a request times out: retrieve the existing build first. Set a per-run cost cap on the agent. The cap bounds effort: the server refuses an effort whose minimum spend exceeds it, whatever cap a caller sends. Choose the cap together with the effort my runs will use and the agent's default mode or depth, using the field names the schema gives them. For application integration, use the current SDK docs at https://grep.ai/developers/api/quickstart or the live REST schema. Select the supported SDK for the project's language and confirm its package name against the package registry before installing; if the documented name is not published, look for the published name in the SDK reference pages, and never install a similarly named package from another author. Write a small integration that follows the codebase's conventions for scripts, linting, and tests. If neither the SDK docs nor the live schema documents the base URL and request shape the SDK needs, do not guess them. Call the documented endpoints with the language's standard HTTP client, and take status values, including terminal states, from the schema or a real response. For the AgentRun create/run contract, configure the deployment-specific AGENTRUN_API_BASE from the SDK docs and use `POST $AGENTRUN_API_BASE/runs` as documented in the quickstart; confirm that path exists in the selected deployment's live schema before writing code against it. Grep v2 research (`POST /api/v2/research`) is documented as its own contract; do not assume the two share a base URL, endpoint, or request shape unless the live schema shows it. Do not use the old `/api/v1/research` endpoint. ## Step 5: Run one task and verify the result Run a representative input within the scope and budget I have authorized. If that authorization is missing, ask before starting billable work and propose the smallest option: one run at the lowest effort. Use the supplied example task as the input when there is one; otherwise use a real input from the project's history rather than an invented one. Send an idempotency key when the docs or the schema support one for run creation (see https://grep.ai/developers/api/idempotency), so a retry cannot start a second billable run; reuse the same key only with an identical body. Save the returned run ID, poll to completion, and inspect the actual output, sources, errors, and usage. Check it against the success criteria; fix configuration or integration errors and report what was verified. When the output makes a claim an outside source can check, verify at least one load-bearing claim against a primary source such as a registry, changelog, advisory database, or official record. When the input is the only source of truth, check the output against the input instead. Well-cited output that passes the schema can still omit a fact, and a miss found this way is a failed test. Read the run's tool calls; few cited sources or searches limited by date are reasons to look closer. Fix the agent's instructions for the cause you found. If you do not run the revised agent, say so. For authentication failures, check the account, scopes, and environment loading. For schema failures, re-read tool discovery or the OpenAPI schema. After a timeout, read the existing run by its ID and keep polling rather than submitting it again; a "continue" operation starts new work on a completed run and is not a way to resume polling. Never present a mock response as a successful live run. ## Step 6: Make the task reusable and improve it Save a project reference to the agent, its purpose, expected inputs, output contract, and an example invocation (for example in `.grep/agents.json`, following the installed skill's format, or the repository's agentify SKILL.md if it is not installed). Keep credentials, sample records, and per-run logs out of this file; it holds identifiers, contracts, and the invocation. If the skill's format has a field for the last verified run ID, fill it; add an example invocation even when the format does not name one. Show me how to rerun it with new inputs from my coding agent or application, and how to read a run by ID after a timeout. When the same procedure recurs in context available to you, suggest or reuse this agent. Skill installation is not an always-on watcher and does not grant access to other conversations. Use completed-run evidence and the installed optimize workflow, when available, to identify repeated deterministic work that can become code or workflow steps. A lookup with an exact answer belongs in code: compute it, pass it to the agent as context, and check the output against it so the integration fails when the agent omits something. Keep model judgment where it is needed. Discover the deployed workflow capabilities, validate a candidate on representative inputs, and compare output quality and measured usage before claiming an improvement. Do not invent a public optimize endpoint or promise that every run costs less. Finish with the installed skills, connection status, agent reference, test result, and exact rerun command or tool call. Explain anything still blocked. Batch execution and scheduling should follow my actual authorization, not be enabled merely by onboarding.Read the full Markdown guide ↗Click the button to copy the full prompt.
Create an agent in minutes
Cost per run
Optimize repeated work
Company researchIllustrative workload · USD




Set up your agent.
Let it get to work.
Describe the job in plain English, or start with a few lines of code. Run the same agent from chat, your application, or an API.
- Describe the job. Give your agent its instructions.
- Run it your way. Use chat, an SDK, the API, or MCP.
- Build on experience. Use completed runs to inform optimization.
Your merchant classification agent is ready.
Classifies merchants, cites the rules, and flags ambiguous cases.
Create an agent for any repetitive task
Give your agent the rules, the evidence to check, and the output you need. Start with one of these examples, then make it yours.
Run it once. Run it 1,000 times. Run it 100,000 times.
Consume fewer tokens as you scale
Business onboarding
Complete these checks before approving a new company.
- 01Verify the business
Confirm registration and identify beneficial owners.
- 02Review the risks
Screen sanctions lists and adverse media. Flag matches.
- 03Document the decision
Write a brief with sources and refer flagged cases for review.
Your merchant classification agent is ready.
Classifies merchants, cites the rules, and flags ambiguous cases.
Describe the job and share your process. Grep uses deep research to build your agent’s skills and instructions.
Illustrative workflow and costs. Actual cost depends on the task and model mix.
Scale at a fraction of the costs
Grep runs on AgentRun. It learns how your work gets done, turns repeated steps into reusable workflows, and brings in deeper reasoning when a case needs it.
The agent builds the workflow
Traces and reusable notes inform a program that can skip work which would not change the outcome.
Decision models handle judgment
Decision models evaluate evidence at each branch. Code handles predictable steps; specialized agents investigate cases that need deeper reasoning.
Exceptions still get attention
Cases the workflow cannot settle return to a stronger agent or the review path required by the procedure.
Claude agent
One 24-record alert · before and after
- Time to complete
- 51 min
- Tool calls
- 826
- Records researched
- 24
AgentRun workflow
One 24-record alert · before and after
- Time to complete
- ~3 min
- Tool calls
- ~30
- Records researched
- 2
Reported AML evaluation: costs are averages across 100 alerts; time, tool calls, and records refer to one 24-record alert. This evaluation used a Claude agent baseline, not Codex. Savings vary by workload. Read the methodology and results ↗
Build your agent workforce.
Create agents around your process. Run them across a list, on a schedule, or together on a larger assignment.
Custom agents. Your process, built in.
Your agent is a filesystem: instructions, context, skills, and tools, connected in one workspace and shaped around your process.
instructions.mdInstructionsYour rules, decisions, and review criteriacontext/ContextPolicies, documents, and worked examplesskills/SkillsReusable methods for each part of the jobtools.jsonToolsConnected APIs, search, and integrationsschema.jsonOutputThe structure every result should follow
One agent. Many inputs.
Run the same agent across companies, documents, or records. Each input gets its own result, using the same instructions and output format.
| Company | Status | Result |
|---|---|---|
| Acme Holdings Ltd | Completed | Report |
| Cedar Technologies | Completed | Report |
| Northstar Trading | Completed | Report |
3 companies researched. 3 reports ready.
Keep watch. Get the update.
Run on a schedule, or let a monitor alert you when something changes.
- What to watch
- Product launches and pricing changes from the companies on my watchlist.
- Check frequency
- Every weekday · 9:00 AM
- Deliver to
- Inbox and email
Multiple agents. One campaign.
Connect agents into a workflow. Each agent passes its findings to the next, so the campaign takes a job from input to finished output.
- Ownership agentVerified owners01
- Risk agentScreening findings02
- Assessment agentFinal recommendation03
Your work. Your format.
Get results ready to read, present, share, or use in another system.
- Report
- Doc
- Deck
- Sheet
- Database
- HTML
- Dashboard
- Audio / Video
- JSON
The managed agent platform
You define the work.
We run the infrastructure.
The operating system for your agents: orchestration, isolated execution, data access, and the controls that keep production runs moving. Define the work; Grep manages the infrastructure behind it.
Explore the developer docsAgent OS
Budgets, credentials, and recovery follow the job.
Batteries included.
100+ permissioned data sources, 250+ research skills, and 25+ built-in expert agents, wired into one fluid system. Compose them yourself, or let an agent pick the right ones for the task. Grep maintains every integration for you.
Your data sources. Already connected.
USPTO
OpenCorporates
ComplyAdvantage
Zillow
OpenSanctions
Trustpilot
Amadeus
Semantic Scholar
PubMed
Crunchbase
Middesk
Finnhub
LEI / GLEIF
Yelp
FRED
GDELT
arXiv
LinkedIn
MarineTraffic
CoinMarketCap
PIPL
ImportGenius
EPO
People Data Labs
Built for production
Reliability goes beyond the model.
Context, tools, evaluation, and recovery all shape the result. The Grep harness gives every run the structure and controls it needs.
The white paper draws on external research, including Clyro’s catalog of 591 documented incidents. These are industry examples, not measurements of failures on Grep.
Context blindness
The agent works from stale, missing, or fabricated information.
Keep the working context
Instructions, skills, evidence, and working notes live in the agent’s workspace. It can read the relevant files again instead of relying on one long conversation.
Rogue actions
It takes actions beyond the permissions or scope the job requires.
Scope each tool call
Tools are granted per job and calls go through an execution gateway. Provider credentials stay behind a proxy, outside the agent’s context.
Silent degradation
Quality declines while outputs still appear plausible.
Evaluate before changing the workflow
Real run traces inform workflow updates. Candidate workflows are tested before activation, and each run keeps a record of the version it used.
Memory corruption
State from another user or session contaminates the current task.
Keep cases separate
Each run has its own identity and workspace. Context and evidence belong to that run, with saved state and tool receipts retained for recovery.
Runaway execution
Loops and repeated calls consume resources without a clear stopping condition.
Bound the work and preserve progress
Turn, tool-call, and cost budgets limit execution. Saved progress and tool receipts let recoverable runs continue without blindly repeating completed work.
Source: Clyro’s 591 documented incidents (2023–2026), cited in Grep’s white paper. This is a non-random sample; the shares shown are not production failure rates.
Run an agent today.
Or deploy an agent workforce in days.
Try it yourself
Sign up, start on credits, and run an agent the same day without talking to sales. Self-serve runs the same harness, sandboxes, and audit trails as enterprise.
Speak to sales
For teams running agents at volume: a working session against your real workflows and constraints. We size the model mix, the cost ceiling, and the audit requirements with you.





