THE AI INSTITUTE / RESEARCH FOR LEADERS
Give the AI Agent a Job Before You Give It Tools
The practical unit of agent deployment is a governed job: one outcome, one owner, scoped authority and a clear way to stop.
The smallest useful unit of agent deployment is not a prompt or model licence but a governed job that can be taught, observed, measured and withdrawn.
Key points
What this paper means for leaders
- Start with a recurring business outcome, not a catalogue of agent features.
- Write the agent's job before connecting tools: trigger, inputs, authority, evidence, review and stop condition.
- Give every agent its own identity and time-bound permissions rather than a person's credentials.
- Test realistic queues, interruptions and exceptions; single-task demonstrations hide operating failure.
- Scale the job only when completed outcomes improve without creating unmanageable review or recovery work.
The capability shift
Agents are starting to carry work, not just answer questions
On 1 September, OpenAI published three cases in which agents had been placed inside recurring work: employee onboarding, account management and developer integration. The cases are supplier-selected and cannot establish representative performance. Their value is the operating pattern. Teams teach a stable process, give the agent persistent context as work changes and let it carry a defined opportunity into action. 1
That is a different proposition from adding a chatbot to every role. A chatbot is primarily an interface. An agent job has a trigger, a queue, access to systems, work that persists after the first exchange and a result that somebody must accept. The executive question is no longer whether the model can produce an impressive answer. It is whether the organisation can assign a piece of work without losing the outcome, authority or recovery path.
The prompt starts an interaction. The job defines what the organisation is prepared to let continue.
The new unit of deployment
A governed job is smaller than a transformation and larger than a task
Broad mandates such as 'use agents in finance' are too vague to manage. Tiny demonstrations are too narrow to reveal operational value. The useful middle unit is one recurring job: reconcile a defined exception queue, prepare one account review, assemble one onboarding path or test one integration against agreed criteria.
A job gives leaders something they can fund and withdraw. It has an accountable owner, a baseline and a boundary. It also lets teams change the model or supplier without rewriting the purpose of the work. The technical system may evolve quickly; the outcome, authority and acceptance test should remain legible to the people running the business.
The job specification
Write nine things before connecting another tool
Define the outcome and its baseline. Name the event that starts the work. List the context the agent may use, the tools it may call and the identity under which it acts. Set the authority it can exercise without review. Specify the evidence it must leave behind, the moments a person must decide, the executive who owns the result and the condition that stops the job.
This is not technical documentation for its own sake. Each element answers a management question. What result are we buying? Who can initiate it? Which information may travel? What can change in the real world? Who can explain the decision? Who carries the exception? When does continued operation become irresponsible or uneconomic?
Purpose
Result, baseline, trigger and accountable owner.
Business designOperating boundary
Context, tools, identity and permission to act.
Control designProof and intervention
Evidence, review points, exception route and stop condition.
Management designThe dependency that changes
Identity and authority move into the workflow
NIST warns that people are already giving agents personal and enterprise credentials. That shortcut weakens accountability and exposes broad, long-lived access across systems. NIST's current direction is to treat agents as first-class entities with their own identifiers, credentials and entitlements, connected to the person or system on whose behalf they operate. 3
For an executive, the principle is simple: an agent should not inherit a person's entire working life. Give the job only the identity and authority it needs, for the time it needs them. Record which person or system delegated that authority. When the job ends, the access should end with it.
Control during execution
Approval before launch is no longer enough
Microsoft's 2026 transparency report describes controls that evaluate agents against policy, place checks at critical points in a workflow and monitor behaviour while the system is running. The important shift is from a static review to an operating capability: see what the agent is doing, test it as conditions change and intervene before a failure becomes the accepted process. 2
Procurement should therefore ask more than whether a supplier has guardrails. Ask whether your organisation can define job-specific policy, observe important actions, pause the work, preserve a record, route an exception and restore service. A control that exists only in a sales presentation is not part of the job.
What remains unproven
Real work is messier than a successful demonstration
A Microsoft Research system designed for concurrent enterprise tasks reports why caution is warranted. Across three agent backends, baseline completion fell as the number of interdependent tasks increased. The researchers found that memory, isolation, planning and learning architecture mattered more than changing the base model alone. The environment was simulated, so the numbers are not a production forecast, but the failure pattern is useful. 4
The newest arXiv batch points in the same direction. Researchers are measuring agent working memory, reconciling process supervision with outcome credit and separating short- from long-horizon industrial tool use. These are preprints with different review status, not purchase recommendations. They show that the unresolved frontier is sustained work under changing conditions—not merely a better first response. 567
Test the Monday morning queue: interruptions, stale context, missing inputs, conflicting priorities and the person who must recover the work.
Where the pattern changes
The job can travel; its operating conditions cannot
North American suppliers and standards bodies are rapidly building the platform, identity and control layers around agents. In Europe, current transparency and general-purpose model obligations now sit beside later high-risk duties, making role, disclosure, logging and supplier information part of deployment planning. 9 The United Kingdom applies existing consumer obligations to agentic services. Other markets may start with fragmented enterprise data, lower connectivity, different languages or fewer assurance resources.
Do not use a US supplier case as a global operating template. Keep the nine-part job specification, then adapt the authority, data path, worker consultation, customer disclosure, language quality, fallback and incident route for each jurisdiction and institution. Australia adds a useful assurance comparison, but it is not a proxy for regional or global practice.
The next 30 days
Commission one agent job—not an agent programme
Choose one recurring workflow with a named owner, measurable stakes and enough volume to learn. Write the nine-part job before selecting more tools. Run it first with narrow authority and visible human review. Test normal work, exception work and recovery after interruption. Measure completed outcomes, quality, cycle time, review load and reversals—not messages, tokens or demonstrations.
Scale only when the job improves the business outcome without creating unmanageable checking or recovery work. Expand authority in explicit stages. If no executive owns the result, if the agent cannot leave a useful record, or if the organisation cannot stop it without losing the workflow, the job is not ready for more autonomy.
The 30-day request: one job, nine fields, narrow authority, realistic exceptions and one explicit stop condition.
Research record
Method and limitations
Method
This briefing compares a current supplier case set with Microsoft's current responsible-AI report, official NIST identity analysis, Microsoft Research on realistic multi-task agent environments and the newest arXiv release batch. Supplier cases, standards guidance, simulated research and preprints retain their different evidentiary status. The public treatment translates the convergence into operating decisions rather than presenting a technical paper digest.
Limitations
The OpenAI cases are supplier-selected and involve three AI-native companies. Microsoft's controls are supplier-described. NIST guidance and agent standards remain in development. CORPGEN uses a simulated environment, and the cited arXiv papers are preprints or conference-labelled submissions rather than broad production evidence. Legal, language, labour, infrastructure and data conditions vary by market.
First published 2 September 2026 · Updated 2 September 2026 ·Research period August 2026 – September 2026 · Research current to 2 September 2026 · Version 1.0 · Suggested citation: The AI Institute, Give the AI Agent a Job Before You Give It Tools (2026).
References
References and source notes
- 01OpenAI, How AI-native companies turn workflows into operating capability ↗
Supplier-selected cases published 1 September 2026; useful for workflow patterns, not representative performance.
- 02Microsoft, Responsible AI in 2026 ↗
Supplier transparency-report summary published 1 September 2026.
- 03NIST, Why Agentic AI Needs a Strong Identity Foundation ↗
Official technical analysis published 27 August 2026.
- 04Microsoft Research, CORPGEN advances AI agents for real work ↗
Research benchmark and simulated multi-task environment; not production ROI.
- 05arXiv, Measure Before You Manage ↗
New preprint in the 31 August batch; working-memory evaluation direction.
- 06arXiv, Reconciling Process Supervision with Outcome-Based Credit ↗
Work-in-progress preprint; does not establish deployment readiness.
- 07arXiv, ATLAS industrial tool-use evaluation ↗
Preprint on short- and long-horizon evaluation; no production proof.
- 08OECD.AI, What agentic AI is and does ↗
Multilateral conceptual work published 3 March 2026.
- 09European Commission, AI Act enforcement framework ↗
Official implementation page updated 24 August 2026; distinguishes current and later duties.
Download
Download the paper.
Get the print-ready PDF and receive future Institute research by email.
This web page is the accessible version of record.