← WorkAI workflow

Once deployed, a PM can launch a live A/B test straight from an approved Figma design

A performance-marketing publisher's experimentation program was gated on engineering capacity: every test, however small, queued behind a build. I rebuilt an internal Figma-to-Optimizely tool from a Claude skill into a Claude Code plugin, defined the layer-naming convention that separates editable fields from structural ones, and took it through PM and engineering review as a workflow the team runs without me.

OutcomesMulti-brand publisher
CS-03 / 06
Engineering involvement to ship a test, once deployed
None
Path from an approved design to a live test
Figma → Optimizely
Roles that can run it, not just me
PM or designer
Architecture rebuilt to clear a structural ceiling
Skill → plugin
Client
Multi-brand publisher
Role
Tool architecture, Figma-side contract design, PM and engineering alignment, documentation
Scope
3 sprints · 2026

Throughput is set by build capacity, not by hypotheses

A testing program's throughput is not set by how many hypotheses the team has. It is set by how many builds engineering can absorb. When every test needs a ticket, the tests that get run are the ones somebody could argue for in a planning meeting, which systematically excludes the small, cheap, high-information ones. The program quietly becomes a queue with opinions attached.

The design problem here was the workflow, not a UI.

Why version one didn't work, and what that diagnosis is worth

The first version was built as a Claude skill. Skills cannot house the supporting files this tool needed to reliably read a Figma file and emit working experiment code, so output quality was a function of luck. The failure was architectural, and so was the fix: rebuild it as a Claude Code plugin, where those files have a home.

That distinction is most of the job in AI tooling. Teams stall at the demo because they keep tuning prompts against a structural ceiling. Knowing which layer you are stuck at is the difference between a workflow that ships and a Slack channel full of screenshots.

The interface decisions are where the design work was

The plugin is the easy half. The contract between a designer's file and the generated code is the design work, and it is where a tool either gets trusted or gets abandoned.

  • A Figma layer naming convention that lets the plugin tell an editable design field from a structural one. That is a design decision, not an engineering one: it determines what a PM can change without breaking the test.
  • Manual versus automatically populated fields, defined with the PM team. Which content a human types and which is filled in for them is a workflow question with a revenue consequence, and getting it wrong means either a stale test or a blocked one.
  • Ownership boundaries for functional elements. Image carousels and similar interactions: generated by the plugin, or configured by the PM after export? Drawing that line explicitly is what stops a tool from being 80% helpful and 100% untrusted.
  • Engineering alignment on the integration approach, settled before rollout rather than discovered during it.

Adoption was a deliverable, not an afterthought

The how-to documentation was written in the same sprint as the rebuild, for one reason: a tool only one person can run is a bottleneck wearing a cape. The measure of an install is whether it still runs when its author is on PTO.

The same principle showed up in the smaller work. Component tickets shipped with a Claude Code dev-handoff file alongside the design, so the path from Figma to production is one artifact instead of an interpretive exercise. One ticket, two deliverables, fewer questions during the build.

Where this stands, precisely

Stakeholder-ready and aligned with engineering. Not confirmed live in production. Every claim on this page is written as once deployed, and it stays that way until it isn't true.

The half-sentence of extra lift from claiming a live rollout is not worth being the one thing on this site that someone could check and disprove.

The install is the product

A demo makes a team optimistic. A documented plugin with a naming convention and a named owner makes them faster on a Tuesday. That is the difference, and this is what it looks like inside a company that already had a testing program and could not feed it.

Next stepCS-03 / 06

Want the same shape of outcome?

Send the problem in plain language. You get a read on whether it’s worth doing, what it would take, and what it would cost — before anyone signs anything.