Start with ready-made AI agents with instructions on how to manage them on the marketplace. Browse the library
Back to blog
Back to blog

Yandex Accelerated QA by 30%: How AI Agents Transformed Testing and Freed Up Hundreds of Hours

https://s3.ascn.ai/blog/4f82984d-ec62-4e54-bfb8-4fc1d58848de.png
ASCN Team
10 July 2026
Build an AI agent for your task
It will handle requests, sort your inbox, compile reports, and follow up with clients. No coding or complex integrations required.
Try for free

Yandex implemented AI agents into its testing process, leading to a 30% acceleration in automated test writing and the daily generation of hundreds of checklists. QA engineers now spend half as much time on routine tasks, focusing on more complex and creative challenges, with overall time savings measured in hundreds of hours per month.

In a large IT company like Yandex, routine in QA processes isn't just a loss of hours; it's a direct hit to product release speed. Endless manual checks, repetitive test cases, monotonous automated test writing—all this distracts qualified engineers from truly important tasks, slows down releases, and increases costs. This doesn't have to be the case, and today, this problem can be solved with AI agents that take over up to 80% of routine operations.

Where Yandex Was Losing Time: Testing as an AI Proving Ground

Yandex, with its abundance of services and continuous development, faced a common problem for large IT companies: routine in testing processes. The QA domain, with its clear structure, repetitive patterns, and high task granularity, had long been considered an ideal candidate for applying generative neural networks. QA engineers spent a significant portion of their time on:

  • Manual creation of checklists and test cases. A process requiring attention to detail but often repetitive for similar functions.
  • Writing typical automated tests. Coding that follows specific patterns and can be automated.
  • Performing routine regression tests. Monotonous verification of existing functionalities after each update.

These tasks, though critically important for product quality, consumed the valuable time of highly qualified specialists, diverting them from more complex analytical and research tasks. The problem was not only in time loss but also in potential employee burnout and slowing down development cycles.

From a "Zoo" of Prototypes to a Scalable Solution

Early MVPs created in various Yandex departments showed that AI agents could successfully generate simple automated tests and decent checklists. However, scaling revealed a significant problem: the quality of AI work sharply dropped when moving beyond narrow scenarios. What worked well for one engineer was unsuitable for a team of 15, let alone a thousand testers across the company.

The rapid emergence of multiple AI prototypes led to a "zoo" of technologies, where each team created its own solutions. This created difficulties with support, standardization, and a lack of common quality metrics. To avoid administrative pressure and maintain team motivation for innovation, Yandex adopted a compromise solution:

  • A central team took responsibility for data infrastructure, developing base models, and measuring quality.
  • Local teams were given freedom to experiment with prompts and agents, adapting them to their specific needs.

A key element was the Test Management System (TMS), integrated with AI tools. TMS became the central point for orchestrating all AI use cases in testing, ensuring seamless operation, control, and standardization of processes.

How AI Agents Were Designed for QA

AI agents were designed as multifunctional assistants, deeply integrated into existing QA processes. Their primary task was to automate routine and repetitive actions, freeing up engineers. Three key areas of work were identified:

  1. Checklist and test case generation. Agents, integrated with TMS, Task Tracker, code repositories, and internal Wiki, gained access to all necessary information about the product and changes. They were required to generate comprehensive and relevant checklists, reducing the time for their creation. To ensure quality, the LLM-As-A-Judge approach was used: one model verifies the results of another against a set of criteria, comparing generated cases with reference ones created by experienced testers. This allowed for a flexible system where specialized models are used for complex scenarios, and a base model, continuously retrained by the central team, is used for others.
  2. E2E automated test generation. Here, agents acted as smart coding assistants. They were required not just to generate code snippets but also to understand context, propose optimal solutions for End-to-End tests, considering the application's architecture.
  3. Manual test execution. This is the most ambitious direction. Agents had to learn functional testing in web and mobile applications, independently interacting with the interface, identifying errors, and documenting them. The goal is to minimize human involvement in regression testing.

Implementation and Results

Implementation proceeded in stages, starting with the least risky but most labor-intensive tasks. The first step was the integration of AI agents for checklist generation. After this process showed stable results, AI assistants were integrated into the automated test writing process. Simultaneously, the development and retraining of agents for manual test execution were underway.

A key success factor was team training. The central team conducted specialized training sessions so that QA engineers could effectively use the new tools and trust them. This approach not only implemented the technology but also changed the work culture, making AI a natural part of daily operations.

Metric Before AI Implementation After AI Implementation
Time for checklist creation Baseline Reduced by 50%
Number of checklists generated Not applicable (manual work) More than 200 daily
Speed of E2E automated test writing Baseline Increased by 30%
Share of active AI usage in teams 0% 30-60%
Accuracy of manual tests by AI agent 0% 45% (target 80%)

Thanks to the implementation of AI agents, Yandex achieved significant results:

  • Time savings. A 50% reduction in checklist creation time and a 30% acceleration in automated test writing freed up hundreds of QA engineers' hours, which can now be focused on more complex tasks such as exploratory testing and incident analysis.
  • Scalability. Daily generation of over 200 checklists allowed maintaining a high pace of development without increasing QA staff.
  • Quality improvement. The LLM-As-A-Judge system ensures a high level of quality for generated artifacts, and continuous agent retraining improves their accuracy.
  • Future prospects. The current accuracy of 45% for manual test execution is already a significant step. The target of 80% will allow running most regression tests without human intervention in the future, shortening the feedback loop and new feature development cycle.

How to Replicate This in Your Business

Yandex's case demonstrates that AI agents can radically change approaches to QA, making processes faster, more efficient, and less costly. If your QA team faces similar routine problems, here's where to start:

  • Identify the most routine tasks. Begin with processes that consume the most time and have clear, repetitive patterns, such as generating basic test cases or writing typical automated tests.
  • Integrate AI agents into existing infrastructure. The less employees have to change their habits and learn new tools, the faster adaptation will be. Utilize existing TMS, repositories, and Wiki.
  • Start with a "soft" implementation. Don't try to automate everything at once. Run pilots for specific tasks, show the team the real benefits, and give them the opportunity to participate in agent development.
  • Develop a culture of trust in AI. Conduct training, explain the principles of agent operation, and demonstrate how they help rather than replace people.
  • Use the LLM-As-A-Judge approach. For critical tasks where quality is paramount, set up a system where one agent verifies the work of another or compares it with human-created benchmarks.

If this case sounds like what's happening in your company, our manager can help: he'll analyze your business and niche for free and point out where an AI agent would bring a real result in your case. Message the manager

MainBlog
Yandex Accelerated QA by 30%: How AI Agents Transformed Testing and Freed Up Hundreds of Hours
ASCN.AI Agent
Exclusive for new users. With your first payment for any subscription plan, you get 2x the subscription duration. Only if you pay today!
By continuing to use our site, you agree to the use of cookies.