Want to explore "what if"?

Death, taxes and assigning tokens

Written by HarryRole: Designer

Wooden artist’s mannequin suspended by strings, flanked by the Figma and Claude logos on a colourful gradient background.

I’ve been playing with the Figma MCP/Claude Code connection for the last few months, and more recently the native Figma Agent. I’ve seen a lot of big claims online about efficiency gains and speed-to-production, but those examples lacked a certain ‘realness’. So rather than add to the noise, I want to do the opposite of a jazzy headline. I'll talk through some real work I’ve been doing with agent skills, and be straight about what 'efficiency gains' actually look like when you refuse to sacrifice quality.

I realised pretty quickly that the biggest barrier to creating meaningful work with these tools is context. You know a lot more about the specific task or client history than you realise, and without that context, agents are shooting in the dark. When it comes to visual exploration, that lack of context can be beneficial. A ‘wildcard concept’, or an approach you hadn’t considered before, can be helpful. But those ‘wildcards’ are starting to look a bit less wild; the same 3 ideas are all over LinkedIn. For that reason, the generative visual design space is something I haven’t been exploring as much. I’m eagerly waiting in the wings to see how that space evolves though.

What I have been exploring is automation. We have a saying here: “If you hate it, automate it” and that has been the backbone for the work I’ve been doing with Figma Agents.

Alphero has developed a strong reputation for our expertise in large-scale design system implementation. We’ve pioneered some powerful ways of driving white-label applications and multi-endpoint design systems using design tokens – long before Figma introduced its native ‘Variables’. While I find this stuff genuinely fascinating (I can’t believe I’m saying that, but it's true), the monotonous task of defining, creating, and assigning tokens to Figma components is the bane of my existence. It's a necessary evil, but if I could wave a magic wand and never have to do it again, I would.

So I did. I created a Figma agent skill that knows everything I know about design systems, and it now does the part of my job that I don’t really want to do. Is it foolproof? No. It still checks in with me every 10 mins to double-check it’s doing a good job, but it’s made a big difference in the way I work. In this article, we’ll discuss a bit of background on how I made it, what it does, and what it means for me as a designer. I hope it helps to clarify what ‘AI efficiency gains’ actually means in the design production space in 2026.

01: How it works

A skill is basically a set of rules for an agent or LLM to follow. It tells them what to do, but often more importantly, what not to do. Agents have a default way of thinking. Their default approach might get you close to what you wanted, but when your task requires a very specific output, it can often leave you wanting more. LLM’s are incredibly capable; you just need to train them, the same way you’d train a new hire.

A skill should break down exactly what you want the agent to do for you; the clearer you are, the more likely you are to get the result you expected. It’s an iterative process – you’ve often got to see where the train falls off the rails, in order to stop it from happening again. Version 1 of this skill took about 10 hours of refinement to get it to a workable state, and more recently, Version 2 took an additional 10 hours of refinement. It is a sizable amount of effort up front, but worthwhile when you see the efficiency gains (and how much it increased my overall job satisfaction).

02: What it does

The skill tells the agent how to build and maintain our design system's foundations directly in Figma, following our team's own conventions rather than generic defaults. It encodes the flexible architecture that we are known for, adapts components across screen sizes and brand themes, audits existing work to ensure no duplication, requests approvals for its approach before building, and checks its own work at the end before seeking final approval from someone in the team. I’ve dumbed it down a bit here for obvious reasons, but it's basically a 5,500 word markdown file that details all the same stuff I would normally do.

Crucially, it also knows when to stop and ask rather than guess. AI hallucination is real, and nipping it in the bud saves us from boatloads of rework and design debt.

03: Efficiencies

The efficiency gains are a difficult thing to quantify. I’ve raced against the agent a couple of times to see who’s quicker, and the agent definitely wins – but it’s not as quick as you might think. I’ve estimated it to be about ~20% faster overall, give or take.

I’d expect the efficiency numbers to increase slightly if it were racing against someone with less experience creating tokens, but that raises an interesting consideration that becomes crucial in defining the true efficiency – that consideration is quality.

There are human checkpoints that are baked in to ensure the agent is staying on course. These manual reviews take time and require a base level of understanding of our systems in order to ensure that the output is of an acceptable standard. The more experienced you are, the quicker this process can be. If you’re less experienced or have less context, you’ll either take longer to approve the agent's work, or be more willing to accept whatever it suggests – which can still occasionally be wrong.

The latter of these two possibilities is what I believe is contributing to the overall ‘averageness’ of a lot of work that's being produced at the moment. We’re blindly trusting these tools to provide the right answer, and not stopping to check that it's giving us what we actually need. Double-checking might slow us down slightly, but it ensures better results.

The other key point to call out is that these agents are not ‘set and forget’. I’m not firing off a prompt and clocking off for the day. I’m just finding other tasks to do in the meantime, while I wait for the agent to ‘ding’ and seek my approval. It’s like having an extra pair of hands doing the stuff I don’t want to do, rather than having a complete digital twin.

04: So, is it worth it?

Yes, but not for the reasons the LinkedIn headlines would have you believe.

If you came here looking for a ‘10x your output’ story, I'm sorry to disappoint. The real number is closer to 20%, and even that comes with a caveat: it only holds if you already know what good looks like. If you strip out the experience of the human who’s moderating the agent, the efficiency increases, but the quality decreases.

That's the part I think gets lost in all the noise. These tools don't remove the need for expertise; they lean on it harder. The 20 hours I spent building and refining this skill were really just me writing down everything I already knew about our design systems; the agent is only as good as the context I gave it. It's not a shortcut around the craft; it's a way of applying it to the bits I'd rather not do by hand.

And that, for me, is the actual win. Not speed, but satisfaction. I've handed off the monotonous, soul-sapping token grind and kept the parts of the job I care about. The agent still dings every ten minutes wanting a pat on the head, and I'm still very much on the hook for the quality of what ships. It's an extra pair of hands, not a digital twin, and I think that's exactly the right expectation to set.

If you hate it, automate it. Just don't expect the automation to have you clocking off at lunchtime.

Written by HarryRole: Designer