Generative ai qe: insights from testing Mobot

Generative AI QE: Insights from testing Mobot

Sakshi Jain
January 28, 2025
4 min read

Table of contents

Generative AI is transforming industries by automating tasks and delivering AI tools, such as our conversational interface, Mobot, to enhance operational efficiency. But these advancements also challenge traditional quality engineering (QE) methodologies.

Unlike conventional software testing, AI models produce dynamic, context-sensitive outputs, requiring a new approach to validation and testing.

At Sumo Logic, we faced similar challenges while testing Mobot. So, how did we streamline best practices for QE when designing a new AI solution? Let’s walk through the strategies we implemented, the lessons we learned, and how they contributed to delivering an optimal AI assistant and log analysis partner for all your log search needs.

What is Mobot?

Mobot is an AI-powered conversational interface that helps you gain insights from logs and resolve issues faster with natural language queries. Whether you need insights or to troubleshoot issues, Mobot converts plain English questions into accurate Sumo Logic queries.

Mobot also provides Explore suggestions, which are recommended queries based on your selected source category, such as AWS WAF. While these features are user-friendly and improve efficiency, they also pose unique testing challenges that require a new testing approach.

Challenges faced while testing Mobot

Testing features of Generative AI models, such as Mobot, differ fundamentally from traditional testing due to the subjective and dynamic nature of their outputs.

Some of the key challenges we faced include:

Strategies we implemented to maintain QE best practices

To address the challenges above, we employed several new testing approaches to adhere to QE best practices.

Reverse prompt engineering

Testing a generative AI model requires diverse and realistic data. We leveraged two key approaches for this:

Performance testing with synthetic data

Testing across more than 200 Sumo Logic apps required simulating large-scale environments. Using an in-house tool, we continuously generated and ingested synthetic data into one organization. While manual testing covered 4–5 apps, this automation helped us evaluate performance at scale and detect bottlenecks effectively.

Sumo-on-Sumo feedback loop

Instead of manually logging issues in Jira, we built a feedback mechanism within the Sumo Logic UI using thumbs-up and thumbs-down buttons. Feedback from these interactions was automatically logged in Sumo Logic dashboards, allowing developers to analyze and address issues efficiently, and ensure rapid feedback, and continuous improvement.

Guardrails for relevance

We implemented strict guardrails to ensure Mobot only responded to Sumo Logic-related queries. Testing these boundaries was critical to prevent irrelevant or misleading responses, safeguarding the product’s usability and customer trust.

Golden data for regression testing

Regression is a key concern with AI models. To mitigate this, we created a golden dataset and integrated it into our automated testing pipeline. Each new prompt or model update was evaluated against this dataset to ensure consistent performance.

Metrics for relevance and accuracy

Collaborating closely with the Sumo Logic AI Model team, we established clear metrics, such as relevance scores and accuracy rates, to monitor the model’s performance and identify issues early. Some additional metrics the team added include:-

Dogfooding and early customer previews for real-world validation

Internal testing and customer previews helped us significantly in gathering diverse feedback to ensure Mobot was performing at its best. We identified edge cases and improved the model based on real-world usage patterns.

Key lessons learned from QE testing

Experience Mobot in action

As Generative AI continues to evolve, so too will our testing strategies. Testing Mobot taught us invaluable lessons about adapting to the unique challenges of Generative AI.

By leveraging techniques like reverse prompt engineering and Sumo-on-Sumo feedback loops, we ensured that Mobot maintained and achieved the standard of accuracy, relevance, and scalability that our customers expect.

Want to try Mobot out for yourself? Try Sumo Logic today with our 30-day free trial.