Agents, tools & MCP

What an agent is made of, the two agent types, how tools run locally or in the cloud, how MCP servers extend them, and how operations are billed.

Agent

An agent is the unit that does the work. Every agent is the combination of four things:

  • a model (the LLM that reasons and decides),
  • instructions (the system prompt that defines its mission and behaviour),
  • tools (the actions it can take), and
  • MCP servers (optional remote tool providers).

There are two kinds of agent:

TypePurpose
pentestAssigned to a pentest phase. These agents carry out reconnaissance, enumeration and analysis.
general / chatFree-form conversational agents for open-ended questions and long-running tasks outside the phase pipeline.

Pentest agents are bound to a specific phase and run as part of an engagement; general/chat agents are driven directly through chat.

Reasoning: effort and thinking

Reasoning models can spend extra compute “thinking” before they answer. Every agent exposes two independent controls over that behaviour, and both default to inheriting the model’s own configuration:

ControlFieldMeaning
EfforteffortHow much reasoning depth the model applies, on a fixed scale.
Thinkingthinking_enabledWhether the model reasons before answering at all.

Effort uses a single canonical scale, ordered from least to most intensive:

none → minimal → low → medium → high → xhigh → max

Higher effort generally means better reasoning at the cost of more latency and tokens. Providers expose reasoning depth differently (OpenAI reasoning_effort, Gemini thinking_level/thinking_budget, Anthropic budget_tokens); the execution backend maps this canonical value to the right provider parameter.

Overrides vs. inheritance

Both fields are per-agent overrides. When they are null the agent inherits the model’s default, so:

  • Effective effort = the agent’s effort, or the model’s default_effort when unset.
  • Effective thinking = the agent’s thinking_enabled, or the model’s reasoning flag when unset.

What the model must support

An agent can only set values its model accepts. Each model advertises its capabilities:

FieldPurpose
supports_effortWhether the model accepts an effort value at all.
effort_valuesThe subset of the scale this model allows.
default_effortThe effort used when the agent doesn’t override it.
reasoningWhether the model can think/reason.
supports_thinking_toggleWhether thinking can be turned off (some models always think).

Setting an effort outside the model’s effort_values, or forcing thinking on a model that can’t do it (or off on a model that can’t disable it), is rejected when the agent is created or updated. If you move an agent to a different model, any override the new model no longer accepts is automatically reset back to that model’s default.

Discovering capabilities

The models catalog (/models in the API, client.agents.models in the SDK, rank models in the CLI) returns these fields so you can pick a valid effort before configuring an agent.

Not to be confused with the streaming thinking event: that is the model’s live reasoning as it works (see Chat), whereas thinking_enabled is the agent-level switch that decides whether the model reasons in the first place.

Tool

A tool is a concrete action an agent can invoke — running a scanner, querying a service, navigating a page. Each tool declares an execution mode that determines where it runs:

execution_modeRuns on
localYour machine, via the CLI — useful for internal targets and software you already have
cloudRank’s infrastructure — nothing to install
bothEither location

The platform ships with more than 35 built-in pentest tools, including nmap, nuclei, gobuster, sqlmap and whois, alongside browser-automation actions. A local tool carries a command template that the CLI fills in and runs on your host.

MCP server

Built-in tools are limited to what’s installed on Rank’s infrastructure. MCP servers (Model Context Protocol) lift that limit by letting an agent connect to a remote server that exposes tools dynamically. The platform never hosts MCP servers itself — you declare your own remote servers and assign them to your agents.

MCP servers connect over one of two transports:

TransportDescription
streamable_httpA remote MCP server reached over HTTP
sseA remote MCP server reached over Server-Sent Events

When a pentest runs, the backend hands the agent’s configuration — its built-in tools plus its assigned MCP servers — to the executor, which connects to each MCP server, discovers its tools, and calls them as the model requests. MCP servers are a feature of the paid tiers.

Operation

Every time an agent gets a response from its model, that response is recorded as an operation. An operation captures the input and output token counts and the computed cost_usd, which is what billing and AI-budget tracking are built on.

Operations are how usage rolls up: each engagement and chat accumulates operations, and their cost is compared against your tier’s monthly AI budget. See Teams & tiers for how budgets work.

How it fits together