Skip to content
All work

Flagship case study

AI-Enabled Automation · UI + API · Agentic Engineering

From test automation
to agentic automation

Playwright UI + API automation powered by Claude, MCP and a self-hosted vector knowledge layer.

“How do you move from writing automation scripts to building a system that can understand, generate, execute, validate and remember automation?”
  • Playwright
  • Claude
  • Playwright MCP
  • Weaviate
  • Docker
  • Jenkins
  • Vector Search

Architecture generalized to protect confidential implementation details.

00Where it started

The question changed

I already had Playwright-based UI and API automation covering smoke and regression testing. The engineering problem was no longer “How do we automate the tests?” It became “How do we change the way automation itself is created?”
  • Executed through Jenkins on scheduled runs
  • Executed across multiple environments
  • Integrated into deployment pipelines
  • Used for both UI and API automation

01The baseline

Playwright
UI + API automation

Hover any stage — or watch the run move through the system.

› Smoke and regression scenarios describing the behaviour to protect.

  • [ Smoke ]
  • [ Regression ]
  • [ UI ]
  • [ API ]
  • [ CI/CD ]
  • [ Multi-environment ]

02The old approach

The first iteration

Initially, we used Markdown instruction files to provide Claude with the context and rules required to generate automation code.

The knowledge was reusable
but not persistent.

The agent could follow instructions, but for new automation scenarios it still needed to rediscover information from the codebase and documentation.

03The engineering problem

The context problem

Every new scenario pulled a wide spread of context into the model. Main engineering question: “How can the agent reuse what it has already learned?”

Context sources

Claude

› Hover a node to see what it contributes.

The compounding cost

04The idea

Give the agent
a memory

I introduced a self-hosted Weaviate vector database as a persistent knowledge layer, running locally through Docker — chosen specifically as agent memory, not as a general-purpose data store.
Context reuse+Cost optimization+Knowledge persistence

05Why vector search?

Retrieve.
Don't rediscover.

If a similar automation scenario has already been understood, the system can retrieve relevant information instead of repeatedly traversing the entire codebase. Genuinely new areas still need exploration — retrieval reduces repeat work, it doesn't replace analysis.

First automation

Next automation

06API automation

API automation

Examples of retrieved context

› Hover a node to see what it contributes.

07UI automation

UI needs more context

API automation can often be generated from API and business context. UI automation additionally requires understanding the actual application interface and its locators — so I introduced Playwright MCP.
  1. Weaviate→
  2. Test intent→
  3. Claude→
  4. Generated steps→
  5. Playwright MCP→
  6. Browser→
  7. UI locators→
  8. Execution

08Generate → execute → validate

The automation loop

The initial browser-driven output is treated as a linear execution flow rather than immediately assuming it is the final framework implementation.

Generate · Execute
Validate · Remember

01 · Test intent

What the engineer wants covered, in plain language.

If fail → analyze → revise → execute again
If pass → generate framework-ready script → execute again
If pass → store relevant knowledge

An engineer stays in the loop to review output — this is assisted, not fully autonomous.

09The knowledge loop

The system remembers

Every validated automation can contribute reusable engineering knowledge for future scenarios.
Claude

Weaviate

API knowledgeUI knowledgeTest contextLocators

↺ feeds back into Weaviate

10The experience

From intent
to automation

A visual demonstration only — press Run to watch the stages. Not connected to any real system.
agent · simulated demo

> Automate the Conda Proxy Repository implementation.

  • Understanding request
  • Retrieving knowledge
  • Matching existing context
  • Retrieving UI information
  • Generating test steps
  • Executing with Playwright MCP
  • Validating
  • Generating framework script
  • Validating again
  • Updating automation memory
Automation ready

11Architecture

The complete system

Hover any component for a short explanation.

↺ back into the knowledge layer

Vector DB

Weaviate knowledge layer

Self-hosted vector database running locally in Docker; stores reusable automation knowledge.

  • LLM
  • RAG / Retrieval
  • MCP
  • Vector DB
  • UI
  • API
  • Execution
  • Validation

12Engineering impact

What changed

Knowledge reuse

Reuse previously understood automation context.

Context optimization

Reduce unnecessary repeated codebase exploration.

UI + API

Support both UI and API automation workflows.

Validation loop

Generated automation is executed and validated before becoming framework-ready.

Persistent automation knowledge

Validated context can be stored for future automation scenarios.

Agentic workflow

Move from instruction-following toward a retrieval, execution, validation and memory loop.

Qualitative outcomes only · [Add measured impact if available]

Automation was no longer just about generating test scripts.

The goal was to build an engineering system that could understand context, reuse knowledge, interact with the application, validate its own output and continuously build reusable automation knowledge.

Playwright × Claude × MCP × Weaviate