Skip to content
All work

Flagship case study

Performance Engineering · Infrastructure · AI Analysis

Performance runs that end with an answer, not just graphs

A Jenkins pipeline that builds, provisions, loads and measures a repository product — then hands the evidence to an AWS Bedrock agent that writes the performance report and, when something is wrong, reads the code to find the root cause.

Performance testing usually ends with a pile of graphs and an engineer squinting at them. I wanted the run to end with an answer.

  • Jenkins
  • Gatling
  • AWS
  • CloudWatch
  • S3
  • AWS Bedrock Agents
  • PostgreSQL
  • React

The story, in ten moves

  1. Build
  2. Provision
  3. Seed
  4. Load
  5. Observe
  6. Archive
  7. Analyse
  8. Report
  9. Debug
  10. Explain

Architecture generalized to protect confidential implementation details.

01The problem

Running the load test was the easy part. Understanding it wasn't.

End-to-end performance testing of a repository product meant building, deploying, seeding data, running load, then manually digging through metrics, logs and reports to figure out what happened — and why.

02The pipeline

Ten Jenkins stages, from JAR to root cause

Replay the run or click any stage. Flip the switch to see how the last two stages only run when the analysis actually finds issues.
Analysis found issues?

Stage 01

Get the JAR files

Pull the exact build under test so every run is tied to a known version of the product.

  • Versioned build artifacts

03Continuous evidence

Every minute of the run is on the record

While Gatling drives load, CloudWatch metrics and logs are collected on a 1-minute interval. The result is a timeline you can correlate — not a single average at the end.

Run dashboard · simulated data

CloudWatch sample · minute 01

p95 latency

49ms

CPU

54%

Heap

53%

Errors

1%

AI performance report · embedded

Summary, degraded endpoints and resource correlation for this run — rendered right next to the graphs.

RCA document · embedded

Appears only when issues were found: suspected cause, affected code paths, supporting evidence.

04AI analysis

An agent that reads the evidence — and then reads the code

Once artifacts land in S3, an AWS Bedrock agent analyses them on the fly and writes the performance report. If it finds problems, it clones the product repository, inspects the relevant code and produces a root-cause analysis document. Significant engineering went into making that debugging stage reliable and conditional.

Analyse

Correlate Gatling results, CloudWatch metrics and logs straight from S3.

Report

Generate a readable performance analysis for every run.

Debug → RCA

Only on issues: clone the repo, trace the code, write the root cause.

05Engineering under constraint

No budget for Grafana. So I built the dashboard platform.

The team wanted Prometheus and Grafana. Budget said no. Instead of dropping observability, I designed a lean results platform of my own.

What we wanted

Prometheus + Grafana for every dashboard and metric.

The constraint

Budget. Running and maintaining that stack wasn't approved.

What I built

My own results platform: a PostgreSQL database on AWS, an ingestion API, and a React web app that shows each run like a Grafana dashboard — with the AI report and RCA embedded.

How a finished run reaches the dashboard

  1. Jenkins last stage
  2. Ingestion API
  3. Read S3 artifacts
  4. Chunk reports
  5. PostgreSQL (AWS)
  6. React dashboard

The final Jenkins stage calls one API. It pulls the run's artifacts, chunks the reports, writes everything to PostgreSQL on AWS — and the web app immediately shows the full run, with the AI report and RCA embedded.

06Ownership

Who built what — honestly

System design

Me

End-to-end architecture of pipeline, analysis and results platform.

Backend & data engineering

Me

Ingestion API, report chunking, PostgreSQL schema.

Pipeline & AI integration

Me

Jenkins stages, S3 hand-off, Bedrock agent flow, conditional RCA.

Web app UI

Built with Claude

React front end produced by prompting, driven by my design and data.

07Principles

What this system is built on

  1. 01A run should end with an answer, not just graphs.
  2. 02Collect evidence continuously — minute by minute.
  3. 03One source of truth per run (S3).
  4. 04Only spend AI effort debugging when there's something to debug.
  5. 05Constraints are design inputs, not blockers.

Dashboard values on this page are simulated · [Add measured impact, e.g. hours saved per run]