

The Amazon AMET Payments team, serving approximately 10 million customers across five countries in the Middle East and North Africa, faced a challenge: generating test cases for each new feature took up to a week of manual labor. After implementing a multi-agent AI solution, this time was reduced to mere hours, significantly improving test coverage quality.
Manual testing is not just a laborious process; it's a source of constant delays and hidden errors that surface only in production. When QA engineers spend weeks on routine test case generation, they not only slow down the release cycle but also lose focus on more complex, creative tasks. This is costly for businesses, as every missed error means potential losses, dissatisfied customers, and reputational risks. But today, there's a solution that automates this routine, leaving humans with control and expertise.
Amazon AMET Payments releases an average of five new features monthly. Each of these required meticulous test case generation, which traditionally took one week of manual effort per project. Quality assurance engineers spent an enormous amount of time analyzing business requirements, design documentation, UI mockups, and historical testing data. This process amounted to the equivalent of one full-time engineer per year dedicated solely to creating test cases.
The problem wasn't just the time cost, but also the quality. Initial attempts with simple AI systems often led to generic phrasing, such as "verify that payment works correctly." However, reality demanded much greater specificity, for example: "verify that when a customer from the UAE selects Cash on Delivery (COD) for an order over 1000 dirhams with a linked credit card, the system displays a COD fee of 11 dirhams and processes the payment through the COD gateway with the order status transitioning to 'awaiting delivery'." Manually detailing such scenarios was exhausting and prone to errors.
The team sought ways to automate this process to reduce time and improve accuracy. Existing test automation tools helped execute already created test cases but did not generate them. Simple AI systems, based on a single agent, could not provide the necessary depth and specificity. They produced generic phrases that required significant human refinement, essentially not solving the problem but merely shifting it from one format to another. It became clear that a more sophisticated approach was needed, one that would mimic the thought processes of an experienced QA engineer capable of breaking down complex tasks into smaller, more manageable parts.
Thus, the idea of a multi-agent system emerged: instead of one "brain" trying to encompass everything, create several specialized agents, each responsible for its part of the process, just as people do in a team.
A key breakthrough in design was a paradigm shift: instead of asking "how should AI think about testing?", the team asked "how do experienced humans think about testing?". This led to a detailed study of the cognitive processes of senior QA specialists, who do not process documents in their entirety, but work in stages, focusing on different aspects.
As a result, a multi-agent AI system called SAARAM (QA Lifecycle App) was developed, comprising several specialized agents, each focusing on a specific aspect of the testing process, mimicking an expert approach:
The system underwent several iterations to overcome the context length limitations of large language models, reduce "hallucinations" (incorrect or fabricated responses), and ensure scalability for handling large volumes of data and complex scenarios.
The implementation of SAARAM began with pilot projects where the system proved its effectiveness on real tasks. Gradually, as the agents learned and improved, the QA team began to trust them with increasingly complex aspects of test case generation. The key was not just creating the technology, but also integrating it into existing workflows to minimize user resistance. Engineers saw that AI agents were not replacing them but freeing them from routine, allowing them to focus on more strategic and complex tasks requiring human intelligence and experience. This made the implementation process organic and successful.
| Metric | Before AI Agent Implementation | After AI Agent Implementation |
|---|---|---|
| Test Case Generation Time | 1 week | Several hours |
| Test Coverage Quality | Baseline, prone to manual errors | Significantly improved, more specific and actionable cases |
| QA Engineer Workload for Generation | Equivalent to 1 full-time engineer per year | Significant reduction, resource liberation |
The implementation of SAARAM not only reduced test case generation time from one week to a few hours but also significantly improved test coverage quality. AI agents generate specific and actionable test cases, reducing the number of errors and omissions that could lead to serious problems in production. Furthermore, the system helps standardize testing approaches and capture the institutional knowledge of experienced testers, making it accessible to the entire team.
The solution is already actively used by the AMET QA team and is planned for expansion to other QA teams within the International Emerging Stores and Payments (IESP) Org, demonstrating its scalability and versatility.
The Amazon AMET Payments case demonstrates that even in complex and critically important areas like QA in financial services, AI agents can bring significant benefits. If your testing team or any other department faces laborious and routine work requiring detail and specificity, this approach can be scaled:
If this case sounds like what's happening in your company, our manager can help: he'll analyze your business and niche for free and point out where an AI agent would bring a real result in your case. Message the manager