Back to all case studies

Case Study Results
by Industry

What changed for 37 clients, in one place. Performance and load testing, functional testing and QA automation, benchmarking and data masking, from a four-day emergency before a launch to a partnership now in its tenth year. To browse and filter the cases themselves, use the case study library.

  • 37 Documented engagements
  • 13 Industries covered
  • 30,000 Peak TPS verified
  • 2008 Working on systems since
  • 25 performance and load testing
  • 8 functional testing and QA automation
  • 3 benchmarking and platform validation
  • 1 data masking

Updated August 11, 2026. Every figure below comes from a delivered engagement.

How Long an
Engagement Takes

The honest answer is that it depends, so here is the range with names
attached to both ends of it.

  • 4 days

    E-commerce launch under a fixed date

    Tynor’s new store was 30x slower than the old one. Diagnosis, fix and proof inside the launch window.

  • 1 week

    One system, one clear question

    PEC needed to know its outage reporting would hold during a storm. One week, run alongside the client and their PaaS provider.

  • 1 month

    Hardware in the loop

    In-flight Wi-Fi tested with real devices rather than simulated users, because the bottleneck lived in the access points.

  • 2 weeks, then 6 months

    Stabilise first, then industrialise

    A corporate bank got its urgent capacity answer in two weeks. The process that keeps the answer true took six months.

  • 7 years

    Testing inside the release train

    FIS Profile ships 11 releases a year. Performance testing runs as a gate on every one of them.

  • 10+ years

    A system we now know better than most

    X5 Retail’s SAP ERP started as an emergency handover before high season and never stopped.

How We Run
a Test

Four steps that stay the same whether the engagement lasts four days
or seven years. The order matters more than the tooling.
  1. Build the load profile from evidence

    A profile assembled from opinions produces a test that passes and a system that fails. We take it from production logs, monitoring and access logs. When there is no production yet, we say which numbers are assumptions and test the sensitivity to them. On a Siebel CRM engagement nobody could hand us a profile at all, so we reconstructed it from the database.

  2. Reproduce the system, not a simplified version

    Emulate what is genuinely out of scope and keep the rest real. For a government ESB we built a REST emulator so a service could be tested without its neighbours. For in-flight Wi-Fi we used physical devices, because the limit was in the access point and no virtual user would have found it.

  3. Test until the numbers stop lying

    One green run is not a result. Moving 6 million users to AWS, eight consecutive runs showed healthy latency; all eight were funded by EBS burst credits that production would exhaust in hours. We keep running until the behaviour is stable and explainable.

  4. Hand back a capability, not a PDF

    The deliverable is the scripts, the profile, the monitoring setup and the knowledge to rerun it. Clients keep testing after we leave: an oil and gas client retained the capability in house, and a bank runs eight quality gates in its own pipeline.

Questions to Ask a
Testing Vendor

How long does a load testing engagement take?

From four days to ten years, and the range is not a hedge. A single question about one system with a known profile takes days: we validated a new e-commerce site in four days and an outage reporting platform in a week. A first full engagement on an enterprise system usually runs four to six weeks. Testing that runs as a release gate is ongoing by design, like the seven-year FIS Profile programme at 11 releases a year.

What should a load testing vendor deliver besides a report?

The test assets, so you can rerun the test without them. That means the scripts, the load profile and how it was derived, the monitoring configuration, and the raw results. A vendor who only hands over a PDF has sold you a snapshot; the value is in being able to repeat the measurement on the next release.

Can you test a system if we have no load profile?

Yes, and this is the normal case. A profile is reconstructed from production logs, database contents, monitoring history or the business plan for a system that does not exist yet. On a Siebel CRM credit approval system we built the profile from the database because nobody in the organisation had one. What matters is that the assumptions behind the profile are written down and testable.

How do you know the test environment reflects production?

You measure the gap instead of assuming it away. Where the test environment is smaller, we scale the load, test the same ratios and state the extrapolation explicitly. Where cloud infrastructure behaves differently over time, we run long enough to exhaust the transient credits: an AWS migration test looked healthy for eight runs until the EBS burst balance ran out.

What does a load test find that functional testing does not?

Behaviour that only appears under concurrency and duration: connection pool exhaustion, lock contention, memory growth, queue build-up, thread starvation, and infrastructure limits like burst credits or network IO. A bank bought 1,000 RPA robots and its system could run 500. Every function worked correctly; the capacity did not exist.

When is load testing not the right thing to buy?

When the question is about correctness rather than capacity, when the architecture is about to change so the results expire before they are used, or when the system has a single obvious bottleneck that a code review would find faster. We say so at scoping. A test that confirms something you already know is an expensive way to feel comfortable.

All 37 Case Studies
by Industry

Thirteen groups, every engagement we have published.

What Load Testing
Will Not Tell You

Three limits worth knowing before you scope an engagement,
with anyone.

  • It does not find functional bugs

    A load test tells you the system slowed down, not that it computed the wrong balance. Those are different questions with different tools. On several engagements we ran functional regression alongside performance work precisely because one does not substitute for the other.

  • A smaller environment gives a direction, not a number

    If the test environment has a quarter of production’s hardware, the result is a ratio and a shape, not a promise about production throughput. It is still worth having. It is not worth quoting as a production figure, and any vendor who does should be asked how they scaled it.

  • An unrepresentative profile invalidates everything

    The most expensive failures we have seen are not test failures. They are tests that passed against the wrong scenario mix, the wrong think time or the wrong data distribution, and gave a team confidence it had not earned. This is why the profile takes the first part of every engagement.

Where Does Your
System Stand?

Tell us what you are running and what you need it to survive. We will tell you what a test would have to look like, and whether it is worth running at all.