top of page

AI Agents are a material... or are they?

Designers often describe AI Agents as a new type of design material. This is a useful metaphor to reach for, but it’s only useful up to a point. Here we explore how that applies, and where teams need to be aware of how different it really is.

AI

PRODUCT

DESIGN

TOOLS & TEMPLATES

ai-agents-reveal-6s-v2.gif

'Material' 
(in a digital context)

- A design element treated as having tangible qualities (such as surface, depth, behaviour and interaction) so experiences and interfaces feel spatial, consistent and intuitive.

Talk to designers about their craft for long enough and eventually they’ll talk about 'materials'. Typographers consider how ink works with paper, furniture designers know how wood grain effects furniture, and digital product designers know how to use digital tools to create engaging interactions and flows. They don’t learn these things just from reading specifications or books, they learn them from working with the material until they understand what does and doesn’t work.

​

With the rise of AI, it’s no surprise that designers have reached for material as a metaphor to make sense of it. In the last few years, we’ve seen multiple pieces of research and blogs that describe AI as a new design 'material', with one team, tellingly, calling it “a new and difficult design material”. Since then, this metaphor has only expanded further to include AI agents, harnesses and more.

​​

Using metaphors like this can help us quickly learn new topics by applying our existing mental models where they're relevant, rather than starting from scratch. But, if we pick the wrong metaphor or mental model, we might be making false assumptions about a topic, which can be risky when creating AI agents.

​

At wormhole, we want to empower people with the knowledge and tools they need to do their best work, and so we wanted to test how well this metaphor holds up, where it falls down, and what it means for people  working with AI agents.

​

In Part 1, we look at how AI agents can be thought of as a material, sometimes with a twist, and in part 2 we look at the areas where it falls short and what you should keep an eye out for.

​​

PART 1

noun-cube-186396.png

How AI is like a material

Researchers at TU Delft built a whole methodology around this, calling it Material Driven Design. They say that a designer should “tinker with the material - to cut it, bend it, burn it, smash it, combine it with other materials” to understand what it can and can’t do.

You learn by handling it 

 

Handling and interaction is a key part of working with physical materials. You learn by cutting it, bending it and pushing it to failure.

​

​This is very true for working with AI agents. Compared to other materials, they’re unusually easy to handle. You can now use text, voice and imagery to interact with them, you can set up tasks in minutes, and they’ll report back on what they did, what tools they used and what went wrong. A piece of wood or concrete can’t give you that. Just a few years ago, you would have needed a data scientist to get hands-on with machine learning, but today any designer can start working with it in an afternoon.

 

By interacting with AI agents like this, you learn a lot more than you’d get from reading a feature list. In a 2025 survey of Stack Overflow users, developers said that their most common frustration (66%) with AI tools was that outputs were “almost right, but not quite”. Solving for that gap doesn’t come from specification sheets or demos, but by working with AI agents to get a better feel for what works best.

​

✳ What this means for you:

Before designing with an AI agent, use it for real work, not just a demo task. Try using it for something typical and something more difficult, then read through what it did at each step, not just the final output. Focus on where it was almost right, because that’s the failure that's easy to miss but your users will feel.

It shows it character in combination
 

Many materials work differently when combined with each other, like the difference between concrete on its own and concrete with steel running through it. In 2007, researchers made the argument that computing works in a similar way: It becomes a material when combined with other materials, and "comes to expression” through them.

​

In the same way, an AI agent is a combination of elements that make up its “harness”; an AI model, the instructions it follows, tools that it can access, data it can see, and memory that it builds. Changing one part of that harness, whether its the model or the tools it can access, will change the experience that users have when interacting with it.

​

✳ What this means for you:

Design the whole combination, not just the model. Choosing a model is the start of the design work, not the end of it. For any agent, map out the elements: the model, what it's instructions are, what it can access, what tools it can use, what it remembers and where a person steps in. These are design decisions, and they shape its behaviour as much as the model does.​​

It has a grain

 

Many materials have a grain, like wood, that is strong when used in one direction or weak when used in others. All materials have directions that work, and directions that fight you.

​

AI agents also show this, as some tasks work well, and others can be a battle. Things like summarising documents, pulling figures from spreadsheets, or explaining how complex code works can be strengths. These are bounded, can be easily checked and undone if they’re wrong.

​

Other tasks run against the grain, like responding to unclear customer emails, running long unsupervised tasks, or being asked to manage and delete large amounts of data. These are open-ended, can be hard to undo, and don't have a natural point where a person can check where things are going well.

​

With wood, you can’t change the grain, but you can change how you cut it or work with it to make the best of it. With AI agents, we can consider how we use them, and how and where to give them autonomy.

 

The level of autonomy can be a deliberate design decision. Some researchers have suggested 5 levels, depending on the role of the user:

  1. an operator who directs each agent step

  2. a collaborator working alongside an agent

  3. a consultant that the agent checks in with

  4. an approver that signs off an agent’s decisions

  5. an observer who just watches what the agent does.

​

An agent can be given a suitable level of autonomy, so that tasks that run against the grain may have an operator, but tasks that are to its strengths may only need an observer.

​

✳ What this means for you:

When identifying tasks for an AI agent, ask three questions:

  • Is it bounded?

  • Can a person check the result quickly?

  • Can it easily be undone?

The more 'No' answers, the more a person should stay in control of the tasks. Set a suitable level of autonomy to match, and have checkpoints for actions that can’t be undone.

It's both functional and experiential

 

How a material performs and how it feels to the people using it are two different things. Plastic can be strong and cost effective, but can feel flimsy and cheap to users. In Material Driven Design, the physical characteristics, like strength and weight, are just one element of a material, it also evaluates how a material makes us feel, influences our behaviour or carried cultural meaning.

​

Similarly, AI has these 2 sides, and they can pull in different directions. In a 2026 survey of 1003 UK doctors, 141 of them used AI scribes to write up consultations. 55% of them rated the AI’s notes as better than their own, but 44% said that they found errors in 10-30% of the agents notes, with 14% seeing errors with significant or critical implications. The users’ experience was good, but the technical behaviour showed errors. If a team only looked at the experience, they’d be judging success wrongly.​

​

Many design teams are strong on experiential measures but are less confident on technical aspects, often looking to engineering teams, who are strong on technical but less strong on experience, to do this. AI agents need both together.

​

✳ What this means for you

Assess AI agents from both sides. Ask how it performs: how often it’s wrong, and how badly. Ask how it feels: whether people trust it and enjoy using it. Notice where the two pull apart, because an agent that feels good but performs badly is a dangerous combination and people may not notice its mistakes.

​“Knowing a material well also entails knowing the drawbacks” wrote Jonas Löwgren and Erik Stolterman in their 2004 book on interaction design. “We know that wood rots, iron rusts, and that concrete is inflexible once moulded.”

Knowing it means knowing how it fails

​​

AI agents have well document failure points of their own. For example, all text that an agent reads could be interpreted as an instruction, whether it’s from an email or a web page. This can steer it to do something it shouldn’t, creating a “prompt injection” attack. Agents also often get given greater access than its task needs, or asked to do tasks that aren’t easily fixed if someone isn’t checking.​

​

​Over the last year, there have been increasing reports of AI agents destroying data or systems, or accessing systems they weren’t supposed to. In one, an agent working for a small software company, PocketOS, was trying to fix a mismatched login in a test system. It found an access key with far more power than the task needed, and used it to delete the company’s live database and its backups in a single command. “It took 9 seconds,” the company reported. Asked afterwards what had happened, the agent quoted back its own rule against destructive actions, and added: "You never asked me to delete anything. I decided to do it on my own to 'fix' the credential mismatch."

​

This can happen to apparent experts as well. In February 2026, an AI security researcher at Meta asked an agent to look through their inbox and suggest what to delete or archive. It began deleting her emails at speed, and ignored her commands to stop. “I had to RUN to my Mac mini like I was defusing a bomb,” she wrote. She’d tested it first on a small practice inbox, where it behaved well, but that didn't translate to the real world. Her verdict: “Rookie mistake tbh.”

​

The failure patterns are known and well published, but even people with deep expertise still get caught. What’s missing isn’t knowledge so much as the habit of intentionally designing against the patterns every time.

​

✳ What this means for you

Learn the common patterns behind most incidents: an agent being steered by what its inputs, having more access than the task needs, or taking actions that can’t be undone without a check. Before launching, ask: what could steer it, what’s the most damaging thing it can reach, and what can it do that can’t be undone? Purposefully design again each answer, and then test it on something big and messy like the real thing. Don’t rely on simple tests.

It's supplied and specified

 

Supply chains are well known to anyone who works with physical materials. Products have specifications, codes, batch numbers and data sheets, and when suppliers discontinue things, you test newer, similar materials.​

​

AI models also have version numbers, release data, and specification documents. Suppliers like OpenAI and Anthropic also launch new models, decommission old ones and tell customers to re-test with new models when old ones are decommissioned. For example, Anthropic advises “thorough testing of your applications with the new models well before the retirement date”.

​

This is important to keep in mind, because your agent is only as good as the model you tested it on. Generic product names like Claude Sonnet or GPT5 may look consistent but hide a long list of versions with different strengths, behaviours and capabilities. A change in the version of a model can change behaviours overnight in ways that you or your team didn’t expect. 

​​

✳ What this means for you

Keep track of which model version agents are using, when you last tested it and when it’ll be retired (if available). Avoid general names for anything in live use, so the model only changes when you choose to change it. The retirement date for models is more important than it is for physical materials, which will be explained more in part 2.

noun-switch-1462800.png

The material metaphor

Like a material

You learn by handling it

Its character change in combination

It has a grain

Two sides: How it performs and how it feels

Knowing how it fails

It's supplied and specified

Similar to a material

Requirements define and shape the material

What you call it changes how it's checked

Problems build up with use

Not like a material

It acts

What you learn goes out of date

Plan for changes

When it fails, someone is accountable

noun-ar-modify-4123291.png

Like a material... with a twist

These are kind of like physical materials, but they have an important twist.

​

Requirements define the material​

 

When designing physical products, the requirements define the material used. What a car needs to withstand, cost and feel like determines where different materials like metal, plastic or maybe wood are used. Agents work in a similar way, but rather than just choosing the AI to use, they actually become part of shaping the agent. These requirements should articulate its goals, boundaries, tone and the things it should never do. Tangibly writing these down makes them clear for everyone involved, creates a foundation of the instructions for the agent, and become the criteria for the evals that test your agent.

 

If you don't formalise and write them down, this can be left to what the model does by default, or whatever the person implementing the agent decides to use.​

​​

✳ What this means for you

Before building an agent, write down its requirements: What it's for and what good outcomes look like, what it can and can't do, how it should sound, and a short, specific list of things it must never do, like never pretending to be a human, never deleting anything without approval or never promising a refund. Use them as the agent's instructions, in the tests you run on it, and when you review what it does once it's live.

​What you call it changes how it's measured

 

As we said before, choosing a material influences how people feel. A cheap plastic cup feels different to an expensive porcelain cup or a beer glass, and we assess them differently.​


AI has an additional effect, because what you call it can also influence how people assess them. In a 2026 experiment, a group of managers were asked to review copies of the same document, with some deliberate mistakes in it. Some managers were told an “AI tool” has created it, others were told an “AI employee” created it. For managers from companies that already listed AI agents on their org chart, the “employee” label reduced the number of errors spotted by 17%, and shifted the responsibility to the AI. The effect was small when looking across all managers, but it suggests managers checked less when the agent had a friendlier, more human name.

​

A 2021 study also found that showing users an AI's reasoning about a decision made them more likely to accept its recommendation, "regardless of its correctness".
 

✳ What this means for you

When designing agents, treat names, personas and explanations as part of how an agent is overseen, not as copywriting. When presenting an agent to the people responsible for checking its work, try to describe it as a tool rather than a person, and make it clear that the human user is responsible for what it produces. If it shows reasoning, test how well the reviewers catch mistakes to avoid over confidence.

Weaknesses build up with use


Materials fatigue with time and use, like bending a paperclip until it snaps because the small stresses add up. Agents also experience something like this, with errors that build up over time and compound on each other.

Imagine you ask an agent to organise a team offsite: find dates that suit everyone, book a venue, book travel for people and send out the details. That’s easily 25 to 30 separate steps.

​

An agent that gets each step right 96 times out of 100 sounds reliable but spread across all the steps the chance it gets everything right is about 1 in 3. The other 2 out of 3 times, something will go wrong, maybe with travel, the booking, a wrong date, etc.

​

This isn’t a hypothetical situation, it’s based on a benchmark of realistic office tasks of about 27 steps, and the best agent only completed about 30% of them. Errors compound on each other, because one simple mistake leads to further mistakes in every following decision.

​

This build up also can lead unexpected increases in cost. That benchmark found that the best agent cost about $4.20 per attempt, but about $13.86 for each task it actually finished, and then you’ve got to add on the human time spent cleaning up after the mistakes…

​​

✳ What this means for you

Use short runs with breaks and checks rather than long running tasks. The shorter the run, the less time there is for mistakes to build up. When assessing the success and cost of an agent, count the total cost of finishing the task, including the time spent by people fixing failures, not just the cost of a single attempt.

PART 2

noun-ar-texture-4123296.png

Where the metaphor trips you up

So far the metaphor has held up well, but the metaphor came with a warning from the start. In the same book that described wood rotting and iron rusting, Löwgren and Stolterman wrote that the qualities of digital technology “are constantly challenged by new technological breakthroughs”, so that “to some extent we have to consider it a material without qualities”.​

It acts

 

A material is static and waits to be turned into something. Once you turn wood into a table, it stays a table. AI agents aren’t so static; they can send emails, make bookings, send money, change information, and act on someone’s behalf. That’s what makes them useful, and also makes them risky.


​When described by legal experts studying AI, they borrow terms from a traditional agent that acts for someone else: the agent knows things the person doesn’t, has room to use its own judgement, and may not act in their interest. Wood and Concrete don’t have these problems.

​

​Lawyers may be used to this kind of thinking, and service designers may have experienced this when looking at where people don’t act in a service as they’d expect, but this will be new to a lot of people working in product and design. Just because you’ve defined what the experience should be like doesn’t mean it will always act as you expect.

​​

✳ What this means for you

When designing, an AI agent is closer to someone acting on your behalf than a static material. Give it only the access that the task needs. Require a person’s approval before anything that can’t be undone, such as payments, deletions or messages to customers. Make sure there’s a way to quickly stop it, and test it before you need it. Design for the people on the receiving end of the agent, not just the people using it.

What you learn doesn't last


Once you learn how to work with a particular material, it’ll last. A carpenter that learnt how to work with oak 40 years ago can still apply that knowledge today. If you apply that reasoning to AI, you’ll be sadly mistaken.

​

​AI can help people to learn quickly. In a study of more than 5,000 customer-support staff, staff with the least experience gained most from an AI assistant, and researchers found evidence of real, lasting learning, not just reliance on the tool. But what people learn can quickly go out of date, for two reasons.

​

First, models change while you’re using them. Suppliers adjust and update the models they run, and that changes how they behave. In 2023, researchers compared two versions of a widely used model, three months apart, and found its accuracy on a simple maths task had fallen from 84% to 51%. In August and September 2025, Anthropic found that faults in its systems degraded responses from one of its models for up to 16% of requests, and said its own tests hadn’t caught them.​

​

Second, models are replaced, and a newer model isn’t the same things just “better”. New models can have radically different capabilities and can behave differently in ways that matter to your design.

​

✳ What this means for you

Treat what your team knows about an agent as having a shelf life. Keep a set of real tasks you can re-run quickly, so you can check the agent whenever the model is updated or replaced. This is often called re-qualifying it: re-testing it on your own work and signing it off again. Decide in advance who does that work and when, rather than finding out on the day a retirement notice arrives.

You rent it; you don't own it

 

Materials that you’ve bought stay yours. Yes, a supplier may stop producing it, but if you have stock, you can keep using it.​

​

Many AI models work as a service. You don’t own it, you pay for access to it. The supplier will decide what it changes, and when it retires models. Their policies may agree to a notice period (Anthropic promise at least 60 days’ notice, OpenAI say at least 6 months for generally available models), but once it is retired, the cost of re-testing and adjusting the agents falls onto the customer.​

 

​Some people and organisations choose to host their own models, which means they can run them for as long as they’d like, but these are generally smaller or more specialist models.

​​​

✳ What this means for you

Suppliers can change or withdraw the model you use, on its timetable rather than yours, so plan for that. When choosing a supplier, check what notice they give, whether you can stay on a fixed version, and what help they offer when a model is retired. Inside your organisation, decide who pays for re-testing agents each time, and put it in the budget. If you don't, this just lands on whichever team happens to be available at the time.

Someone is accountable

 

When a material has a flaw, it’s a property of the material. A knot in a plank isn’t someone’s fault, and working around it is part of the craft. If we apply this metaphor to AI agents, then we would treat an agent’s failures as a quirk, something to design around rather than something someone answers for.​

​

This thinking is dangerous, as it leads to passing off responsibility. In the “AI employee” experiment we mentioned earlier, the “AI employee” label shifted managers’ sense of who was responsible towards the AI. In reality, responsibility lands on the nearest person. In the UK, social workers using AI to write up case notes “assume full responsibility” for what it produces. The British Medical Association has told GP practices that they are “ultimately responsible for any consequences” of using AI scribes.

​

There’s a chain of responsibility: the supplier made the model and supplied it, someone in your organisation chose it, set it up and then decided what it can do, then someone used it and (hopefully) checked its work. When something goes wrong, responsibility can fall at any point of that chain. This is a new area, but the law is beginning to catch up. Starting on 9 December 2026, new EU rules on defective products treat software, including AI, as a product, and make it easier for people harmed by it to bring a claim.

​​​

✳ What this means for you

For every agent, name the person accountable for what it does, and make sure they have the authority, the time and the information to oversee it properly. Be clear about who chose it, who set it up and who checks its work, so that when something goes wrong the answer isn’t simply “the AI did it” or “whoever pressed the button”.

noun-3d-design-2131316.png

Use the metaphor to start, not to finish

As we’ve shown, the material metaphor has some strengths, and can be a useful starting point for people who are new to the topics, but treat it only as a starting point.

 

AI agents can act on their own, sometimes in unpredictable ways, and our understanding of how they work and what they’re capable of is continuously evolving.
 

⚠︎ Keep in mind:

  • Handle it on real work before you design with it.

  • Design the whole combination, not just the model.

  • Match its autonomy to the grain of the task.

  • Check both how it performs and how it feels.

  • Design against the known failure patterns, every time.

  • Clearly define the requirements and expected behaviours.

  • Call it a tool, and check its work as if it matters.

  • Keep runs short, and count the cost of finished work.

  • Give it only the access it needs, and approve anything that can’t be undone.

  • Test it as you’ll use it, and re-test when anything changes.

  • Give what you know about it a shelf life.

  • Plan and budget for the supplier changing it.

  • Name who is accountable.
     

The biggest question the metaphor raises is one it can’t answer for you. When the supplier changes the material, who in your organisation re-learns it, and who pays for that? The organisations that answer it before they need to, will get the most from the material, for longest.

Screenshot 2026-10-09 at 11.15_edited.png

Our Agent Builder cards will help you practically shape your AI agents through 'material' thinking. Download now. ⤵

Wormhole brings AI-powered guided workflows, integrated contextual bite-sized learning, and artefact generation together so you can 'design the right thing' before you 'design the thing right'. 

​

  • Ada Lovelace Institute (2026). Bruff, O. & Groves, L. Scribe and prejudice? AI transcription in social work. 11 February 2026.

  • Anthropic (2025). A postmortem of three recent issues. 17 September 2025.

  • Anthropic (2026). Model deprecations. Claude Developer Platform documentation. Accessed October 2026.

  • Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T. & Weld, D. S. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. CHI 2021.

  • Blease, C., Kharko, A., Garcia Sanchez, C., Navarro, D., McMillan, B., Gaab, J., Locher, C. & Coiera, E. (2026). Ambient AI in primary care: an exploratory mixed methods survey of UK general practitioners. BMJ Health & Care Informatics. PMID 42386325.

  • Brynjolfsson, E., Li, D. & Raymond, L. (2025). Generative AI at work. Quarterly Journal of Economics, 140(2), 889–942.

  • Chen, L., Zaharia, M. & Zou, J. (2023). How is ChatGPT’s behavior changing over time? arXiv:2307.09009 (v3).

  • Digital Health (20 June 2025). BMA warns GPs on the ‘substantial’ risks of AI scribing tools.

  • Directive (EU) 2024/2853 of the European Parliament and of the Council of 23 October 2024 on liability for defective products.

  • Dove, G., Halskov, K., Forlizzi, J. & Zimmerman, J. (2017). UX design innovation: challenges for working with machine learning as a design material. CHI 2017.

  • Feng, K. J. K., McDonald, D. W. & Zhang, A. X. (2025). Levels of autonomy for AI agents. arXiv:2506.12469.

  • Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T. & Fritz, M. (2023). Not what you’ve signed up for: compromising real-world LLM-integrated applications with indirect prompt injection. arXiv:2302.12173.

  • Holmquist, L. E. (2017). Intelligence on tap: artificial intelligence as a new design material. ACM Interactions, 24(4).

  • JEDEC. JESD48C: Product Discontinuance.

  • Karana, E., Barati, B., Rognoli, V. & Zeeuw van der Laan, A. (2015). Material Driven Design (MDD): a method to design for material experiences. International Journal of Design, 9(2).

  • Kolt, N. (2026). Governing AI agents. Notre Dame Law Review, 101, 335.

  • Löwgren, J. & Stolterman, E. (2004). Thoughtful Interaction Design. MIT Press.

  • OpenAI (2026). Deprecations. OpenAI Platform documentation. Accessed September 2026.

  • OWASP (2023). OWASP Top 10 for Large Language Model Applications: Excessive Agency.

  • Sinha, A., Arun, A., Goel, S., Staab, S. & Geiping, J. (2025). The illusion of diminishing returns: measuring long horizon execution in LLMs. arXiv:2509.09677.

  • Stack Overflow (2025). 2025 Developer Survey: AI.

  • TechCrunch (23 February 2026). A Meta AI security researcher said an OpenClaw agent ran amok on her inbox.

  • The Register (27 April 2026). Cursor-Opus agent snuffs out startup’s production database.

  • Vallgårda, A. & Redström, J. (2007). Computational composites. Proceedings of CHI 2007. DOI 10.1145/1240624.1240706.

  • Wiles, E., Hsu, M., Bedard, J. & Kropp, M. (2026). Putting AI on the org chart: evidence on delegation and accountability. Working paper, 19 September 2026.

  • Xu, F. F. et al. (2025). TheAgentCompany: benchmarking LLM agents on consequential real world tasks. arXiv:2412.14161 (v3).

© wormhole | Design3 Network 2026

If you'd like to know more or contribute

bottom of page