Start with ready-made AI agents with instructions on how to manage them on the marketplace. Browse the library
Back to blog
Back to blog

Amazon AMET Payments Reduced Test Case Generation from Weeks to Hours: How an AI Agent Improved QA Quality

https://s3.ascn.ai/blog/d64fe356-7dc0-49ba-9a74-efcbdb815e63.png
ASCN Team
28 June 2026
Build an AI agent for your task
It will handle requests, sort your inbox, compile reports, and follow up with clients. No coding or complex integrations required.
Try for free

The Amazon AMET Payments team, serving approximately 10 million customers across five countries in the Middle East and North Africa, faced a challenge: generating test cases for each new feature took up to a week of manual labor. After implementing a multi-agent AI solution, this time was reduced to mere hours, significantly improving test coverage quality.

Manual testing is not just a laborious process; it's a source of constant delays and hidden errors that surface only in production. When QA engineers spend weeks on routine test case generation, they not only slow down the release cycle but also lose focus on more complex, creative tasks. This is costly for businesses, as every missed error means potential losses, dissatisfied customers, and reputational risks. But today, there's a solution that automates this routine, leaving humans with control and expertise.

The Reality of the Problem: A Week of Manual Work for Each Test Case

Amazon AMET Payments releases an average of five new features monthly. Each of these required meticulous test case generation, which traditionally took one week of manual effort per project. Quality assurance engineers spent an enormous amount of time analyzing business requirements, design documentation, UI mockups, and historical testing data. This process amounted to the equivalent of one full-time engineer per year dedicated solely to creating test cases.

The problem wasn't just the time cost, but also the quality. Initial attempts with simple AI systems often led to generic phrasing, such as "verify that payment works correctly." However, reality demanded much greater specificity, for example: "verify that when a customer from the UAE selects Cash on Delivery (COD) for an order over 1000 dirhams with a linked credit card, the system displays a COD fee of 11 dirhams and processes the payment through the COD gateway with the order status transitioning to 'awaiting delivery'." Manually detailing such scenarios was exhausting and prone to errors.

The Journey to AI Agent: Why Existing Solutions Fell Short

The team sought ways to automate this process to reduce time and improve accuracy. Existing test automation tools helped execute already created test cases but did not generate them. Simple AI systems, based on a single agent, could not provide the necessary depth and specificity. They produced generic phrases that required significant human refinement, essentially not solving the problem but merely shifting it from one format to another. It became clear that a more sophisticated approach was needed, one that would mimic the thought processes of an experienced QA engineer capable of breaking down complex tasks into smaller, more manageable parts.

Thus, the idea of a multi-agent system emerged: instead of one "brain" trying to encompass everything, create several specialized agents, each responsible for its part of the process, just as people do in a team.

How the AI Agent Was Designed: A Human-Centered Approach

A key breakthrough in design was a paradigm shift: instead of asking "how should AI think about testing?", the team asked "how do experienced humans think about testing?". This led to a detailed study of the cognitive processes of senior QA specialists, who do not process documents in their entirety, but work in stages, focusing on different aspects.

As a result, a multi-agent AI system called SAARAM (QA Lifecycle App) was developed, comprising several specialized agents, each focusing on a specific aspect of the testing process, mimicking an expert approach:

  • Customer Segmentation Agent. This agent is responsible for identifying different user segments based on their characteristics and behavior, creating decision matrices, and developing end-to-end scenarios that consider the specifics of each segment.
  • User Journey Mapping Agent. Its task is to build flow diagrams and user action sequences, generate end-to-end scenarios and detailed testing steps, simulating how a user interacts with the system.
  • Segment and Path Coverage Agent. This agent combines data obtained from previous agents to create detailed, segment-specific analyses, ensuring comprehensive test scenario coverage.
  • State Transition Agent. Analyzes various product state points in user journeys, creating state diagrams and generating related test scenarios to ensure all possible state transitions are correctly handled.

The system underwent several iterations to overcome the context length limitations of large language models, reduce "hallucinations" (incorrect or fabricated responses), and ensure scalability for handling large volumes of data and complex scenarios.

Implementation: From Pilot to Standard

The implementation of SAARAM began with pilot projects where the system proved its effectiveness on real tasks. Gradually, as the agents learned and improved, the QA team began to trust them with increasingly complex aspects of test case generation. The key was not just creating the technology, but also integrating it into existing workflows to minimize user resistance. Engineers saw that AI agents were not replacing them but freeing them from routine, allowing them to focus on more strategic and complex tasks requiring human intelligence and experience. This made the implementation process organic and successful.

Results

Metric Before AI Agent Implementation After AI Agent Implementation
Test Case Generation Time 1 week Several hours
Test Coverage Quality Baseline, prone to manual errors Significantly improved, more specific and actionable cases
QA Engineer Workload for Generation Equivalent to 1 full-time engineer per year Significant reduction, resource liberation

The implementation of SAARAM not only reduced test case generation time from one week to a few hours but also significantly improved test coverage quality. AI agents generate specific and actionable test cases, reducing the number of errors and omissions that could lead to serious problems in production. Furthermore, the system helps standardize testing approaches and capture the institutional knowledge of experienced testers, making it accessible to the entire team.

The solution is already actively used by the AMET QA team and is planned for expansion to other QA teams within the International Emerging Stores and Payments (IESP) Org, demonstrating its scalability and versatility.

How to Implement This in Your Business

The Amazon AMET Payments case demonstrates that even in complex and critically important areas like QA in financial services, AI agents can bring significant benefits. If your testing team or any other department faces laborious and routine work requiring detail and specificity, this approach can be scaled:

  • Identify routine, repetitive tasks. Find processes where employees spend a lot of time analyzing documents, gathering information, and generating similar solutions or cases. This can include not only QA but also legal departments, customer support, and onboarding.
  • Break down complex processes into stages. Instead of trying to automate everything at once with a single agent, identify key stages, as Amazon did with customer segmentation, journey mapping, and state analysis.
  • Create specialized agents. For each stage, design a separate AI agent that will perform a narrow but deep function. This will help avoid "hallucinations" and increase accuracy.
  • Integrate agents into your workflow. Ensure that AI agents easily integrate into existing tools and systems to minimize resistance and ensure rapid adoption by employees.
  • Start with a pilot and scale. Launch the system on a small, controlled segment, gather feedback, and gradually expand functionality and scope, proving value at each stage.

If this case sounds like what's happening in your company, our manager can help: he'll analyze your business and niche for free and point out where an AI agent would bring a real result in your case. Message the manager

MainBlog
Amazon AMET Payments Reduced Test Case Generation from Weeks to Hours: How an AI Agent Improved QA Quality
By continuing to use our site, you agree to the use of cookies.