AI agents

Agents that work. Not demos.

A useful agent runs tasks end to end in your systems, under supervision: we design, evaluate and ship them — with the guardrails that come with.

Governed by design

What makes an agent trustworthy

Trust comes from engineering, not from promises.

Explicit action scope

Each agent has a closed list of allowed actions, caps and reachable systems — everything else is denied by default.

Human supervision

Actions that commit the organisation — payment, external send, contract change — wait for a human approval in the workflow.

Full logging

Log the requests, approvals, tool calls and outcomes needed to review the workflow, without exposing secrets or unnecessary personal data. Logs do not reveal a model's full internal reasoning.

Stop & reversibility

Define how to suspend the agent and handle incidents. For actions that cannot be undone, use preventive controls and explicit approval before execution.

FAQ

Frequently asked

What's the difference between a chatbot and an AI agent?

A chatbot answers questions; an agent runs tasks: it plans, acts in your systems through secure connectors, checks its result and reports back. The difference is in the engineering — action scopes, supervision, logging — far more than in the model used.

Do we have to choose between GPT agents and Claude agents?

No. The right choice depends on the task: we evaluate both model families on your real cases and keep the best performer — or a hybrid architecture that combines them. Hunter BI partners with both networks; the recommendation stays guided by your measurements.

How do you stop an agent from making a costly mistake?

With technical guardrails, not instructions: a closed list of allowed actions, caps, mandatory human approval for operations that commit the organisation, full logging and an immediate stop procedure. The agent is tested against failure scenarios before any production.

How long to get an agent into production?

The scope, interfaces, test material and approval requirements determine the schedule. We agree the pilot and production gates after scoping, with reusable components considered where they have already been validated.

Choose a bounded workflow, not unlimited autonomy

A first agent should carry out a task that a responsible person can describe and verify. Define its inputs, the systems it may reach, the operations it may propose and the point at which it must stop. A complex multi-agent architecture is not a prerequisite for a useful result.

Possible starting points include gathering information for a support reply, preparing an ERP summary or helping a developer investigate a defect. These are scenarios to qualify, not promises that every organisation can deploy them unchanged. The available data and interfaces determine the feasible scope.

Separate reading, preparation and execution

Reading a document, drafting a record and sending it are different permissions. For an action that commits the organisation, show the target and content to the authorised approver before execution. A subsequent change in the underlying record may invalidate the approval.

The agent should report a refused operation, missing context or uncertain result explicitly. Tool logs can record requests and outcomes; they do not expose the model's full internal reasoning. Some external actions cannot be undone, so prevention, approval and incident handling matter more than a generic promise of reversibility.

Measure the full task, including human review

Build an evaluation set with expected results and representative errors before the pilot. Include invalid references, denied access, unavailable systems and attempts to expand the task through retrieved instructions. Check the business result, not simply whether a tool returned a success message.

Compare the time to an accepted outcome with the existing workflow. Record review and correction effort, meaningful errors and running costs. The rollout decision should be limited to the tested task and conditions. The operating plan names incident owners, stop controls and the checks required after changes.

Put a governed agent to work?

Explicit scope, human approval, full logging: let's frame an agent that earns your trust — from pilot to production.