Agents, tools & MCP
What an agent is made of, the two agent types, how tools run locally or in the cloud, how MCP servers extend them, and how operations are billed.
Agent
An agent is the unit that does the work. Every agent is the combination of four things:
- a model (the LLM that reasons and decides),
- instructions (the system prompt that defines its mission and behaviour),
- tools (the actions it can take), and
- MCP servers (optional remote tool providers).
There are two kinds of agent:
| Type | Purpose |
|---|---|
pentest | Assigned to a pentest phase. These agents carry out reconnaissance, enumeration and analysis. |
general / chat | Free-form conversational agents for open-ended questions and long-running tasks outside the phase pipeline. |
Pentest agents are bound to a specific phase and run as part of an engagement; general/chat agents are driven directly through chat.
Reasoning: effort and thinking
Reasoning models can spend extra compute “thinking” before they answer. Every agent exposes two independent controls over that behaviour, and both default to inheriting the model’s own configuration:
| Control | Field | Meaning |
|---|---|---|
| Effort | effort | How much reasoning depth the model applies, on a fixed scale. |
| Thinking | thinking_enabled | Whether the model reasons before answering at all. |
Effort uses a single canonical scale, ordered from least to most intensive:
none → minimal → low → medium → high → xhigh → max
Higher effort generally means better reasoning at the cost of more latency and tokens. Providers expose reasoning depth differently (OpenAI reasoning_effort, Gemini thinking_level/thinking_budget, Anthropic budget_tokens); the execution backend maps this canonical value to the right provider parameter.
Overrides vs. inheritance
Both fields are per-agent overrides. When they are null the agent inherits the model’s default, so:
- Effective effort = the agent’s
effort, or the model’sdefault_effortwhen unset. - Effective thinking = the agent’s
thinking_enabled, or the model’sreasoningflag when unset.
What the model must support
An agent can only set values its model accepts. Each model advertises its capabilities:
| Field | Purpose |
|---|---|
supports_effort | Whether the model accepts an effort value at all. |
effort_values | The subset of the scale this model allows. |
default_effort | The effort used when the agent doesn’t override it. |
reasoning | Whether the model can think/reason. |
supports_thinking_toggle | Whether thinking can be turned off (some models always think). |
Setting an effort outside the model’s effort_values, or forcing thinking on a model that can’t do it (or off on a model that can’t disable it), is rejected when the agent is created or updated. If you move an agent to a different model, any override the new model no longer accepts is automatically reset back to that model’s default.
The models catalog (/models in the API, client.agents.models in the SDK, rank models in the CLI) returns these fields so you can pick a valid effort before configuring an agent.
Not to be confused with the streaming
thinkingevent: that is the model’s live reasoning as it works (see Chat), whereasthinking_enabledis the agent-level switch that decides whether the model reasons in the first place.
Tool
A tool is a concrete action an agent can invoke — running a scanner, querying a service, navigating a page. Each tool declares an execution mode that determines where it runs:
execution_mode | Runs on |
|---|---|
local | Your machine, via the CLI — useful for internal targets and software you already have |
cloud | Rank’s infrastructure — nothing to install |
both | Either location |
The platform ships with more than 35 built-in pentest tools, including nmap, nuclei, gobuster, sqlmap and whois, alongside browser-automation actions. A local tool carries a command template that the CLI fills in and runs on your host.
MCP server
Built-in tools are limited to what’s installed on Rank’s infrastructure. MCP servers (Model Context Protocol) lift that limit by letting an agent connect to a remote server that exposes tools dynamically. The platform never hosts MCP servers itself — you declare your own remote servers and assign them to your agents.
MCP servers connect over one of two transports:
| Transport | Description |
|---|---|
streamable_http | A remote MCP server reached over HTTP |
sse | A remote MCP server reached over Server-Sent Events |
When a pentest runs, the backend hands the agent’s configuration — its built-in tools plus its assigned MCP servers — to the executor, which connects to each MCP server, discovers its tools, and calls them as the model requests. MCP servers are a feature of the paid tiers.
Operation
Every time an agent gets a response from its model, that response is recorded as an operation. An operation captures the input and output token counts and the computed cost_usd, which is what billing and AI-budget tracking are built on.
Operations are how usage rolls up: each engagement and chat accumulates operations, and their cost is compared against your tier’s monthly AI budget. See Teams & tiers for how budgets work.