logo
Get In Toucharrow icon
Get In Toucharrow icon
logo

A team of 400+ experts delivering comprehensive end-to-end solutions combining power, functionality, and reliability with flexibility, agility, and usability.

maillogosales@thinksys.comlogo+1-408-837-5515

Quality Engineering

  • Software Testing Services
  • QA Automation Services
  • Playwright Automation Testing
  • Performance Testing
  • Mobile App Testing
  • Cloud Testing

Software Development

  • Custom Software Development
  • SaaS Application Development
  • Mobile App Development

Specialized Testing

  • AI Application Testing
  • Blockchain Testing
  • Security Testing
  • API Testing

Explore

  • All Servicesarrow icon
Clients LoveClutchZero Trust

Ask AI About Us

OpenAIOpenAIPerplexityPerplexityGrokGrokClaude.aiClaude.ai

Follow Us

iconiconiconiconicon

© 2026 ThinkSys Inc. All rights reserved.

  • Privacy Policy
  • Terms and Conditions
Loading blog details...

How ThinkSys Ensured 90% More Stable Releases for Cosmon’s AI-Powered Engineering Applications

Summarize With:
Open AIOpen AIPerplexityPerplexityGrokGrokClaude.aiClaude.ai
  1. homeiconhomeicon
  2. Blogshomeicon
  3. How ThinkSys Ensured 90% More Stable Releases for Cosmon’s AI-Powered Engineering Applications

Cosmon develops AI-powered software that supports engineering workflows across desktop applications, web interfaces, and CAD/CAE platforms.

The company did not have a dedicated QA team and approached ThinkSys to assess the application, identify defects, and understand whether any issues could affect key business workflows.

Testing this type of software required more than conventional functional testing. The QA process also needed to verify Windows desktop compatibility, third-party engineering integrations, AI-generated responses, agent actions, and the final results produced inside connected engineering applications.

cosmon-result.png

ThinkSys introduced a structured quality engineering approach that combined manual testing, targeted automation, compatibility validation, AI/LLM evaluation, and defect management.

This case study explains the challenges Cosmon faced, the QA approach implemented, the testing process, and the results achieved.

Meet Cosmon

Cosmon is an AI-focused software company developing Nexus, a platform designed to support and automate engineering workflows involving CAD, CAE, and other engineering applications.

The QA engagement focused primarily on two applications:

  • Nexus Connector V1
  • Nexus Desktop V2

For Nexus Connector V1, testing covered browser workflows, the connector application, and integrations with engineering platforms including SolidWorks, Ansys, Abaqus, and COMSOL.

For Nexus Desktop V2, ThinkSys continued functional and integration testing while building an automation framework designed to improve regression coverage and execution efficiency.

The wider compatibility scope covers around 6 CAD/CAE applications, approximately 12 application versions, and 4 operating system configurations.

The Challenge

Cosmon’s product operates across several technologies and application environments. This created a wider testing scope than a standard web or desktop application.

The main challenges included:

  • No dedicated QA team: Cosmon needed an external team to independently assess the product and identify issues that had not been caught during development.
  • Complex multi-application compatibility: Nexus needed to work across different CAD/CAE applications, software versions, operating systems, dependencies, configurations, and application states.
  • Windows desktop testing: Testing had to account for permissions, file operations, dependencies, application launches, system configurations, and interactions between multiple desktop applications.
  • CAD/CAE integration validation: Confirming that an integration connected successfully was not enough. The team also needed to validate model operations, geometry changes, file handling, and the final result inside the engineering application.
  • AI and LLM variability: The same prompt could return different valid responses. Exact text matching was therefore not enough to determine whether a test had passed.
  • End-to-end agent validation: A correct AI response did not always mean that the correct action was performed. The full workflow had to be validated from prompt to final application state.
  • Growing regression effort: As more applications, versions, and workflows were added, the regression scope became harder to manage manually.

The Proposed Solution

ThinkSys recommended a hybrid QA strategy that combined manual testing with targeted automation.

The approach focused on the following areas:

  • Structured test coverage: Build a test matrix covering operating systems, CAD/CAE applications, versions, configurations, integrations, AI models, workflows, and negative scenarios.
  • Risk-based defect management: Categorize defects based on severity and business impact so critical issues could be addressed first.
  • Manual testing where judgment was required: Keep exploratory testing, new functionality, usability, compatibility, and complex AI-driven scenarios manual.
  • Browser and workflow automation: Use Playwright with Python for stable and repeatable browser-based regression scenarios.
  • Windows desktop automation: Use Pywinauto and PyWin32 for desktop application control, UI interactions, file operations, application state checks, and connector workflows.
  • AI/LLM validation: Use Promptfoo to evaluate expected behavior, acceptable response variations, assertions, model comparisons, and negative scenarios.
  • Repeatable execution and reporting: Maintain automation in GitHub, execute workflows through GitHub Actions, and use Pytest reports for test results.
  • Complete workflow validation: Test the full path from the initial prompt to the final action and expected result inside the connected engineering application.

Results

The engagement gave Cosmon a clearer view of product quality and helped the team address high-risk issues quickly.

During the initial assessment, ThinkSys identified around 15–20 defects. Approximately 7–8 of these were business-critical bugs that could affect important workflows or expected application behavior.

The defects were then prioritized based on severity and business impact.

Within approximately one month, the team helped reduce the identified defect count by almost 50% through focused testing, fixes, retesting, and regression cycles.

The QA process also improved visibility for the Cosmon team. A shared dashboard allowed the client to see the status of defects, testing progress, fixes, retesting, and closures in one place.

Other outcomes included:

  • Improved coverage across CAD/CAE applications and operating environments
  • Better prioritization of critical issues
  • Stronger regression coverage
  • More consistent validation of AI-driven workflows
  • Better visibility into the relationship between AI responses and actual application behavior
  • A QA framework that could expand as the product added more integrations and versions

With all these efforts, Cosmon achieved 90% more stable releases and confidence in the application. 

Implementation Steps

Step 1: Assessing Requirements and Testing Risks

ThinkSys started by understanding the product architecture, critical workflows, supported environments, integrations, and areas with the highest business risk.

The first phase focused on exploratory and functional testing to establish the current quality baseline.

This helped the team identify high-risk areas and decide which scenarios required manual investigation and which could later be automated.

Step 2: Building a Structured Test Path Matrix

ThinkSys created a structured test matrix to organize coverage across supported combinations of environments and workflows.

The matrix included:

  • Operating systems
  • CAD/CAE applications
  • Application versions
  • Environment configurations
  • Dependencies and permissions
  • Integration scenarios
  • Agent workflows
  • AI/LLM models
  • Functional workflows
  • Negative and edge-case scenarios

This gave the QA team a consistent way to identify coverage gaps and prioritize the combinations carrying the greatest risk.

Step 3: Establishing a Manual Testing Baseline

Before automating critical scenarios, the team validated workflows manually to establish expected application behavior.

Manual testing covered:

  • New functionality
  • Exploratory scenarios
  • Compatibility
  • CAD/CAE integrations
  • AI-driven workflows
  • Edge cases
  • Usability

This baseline also helped define the acceptance criteria used later in automated testing and AI evaluations.

Step 4: Developing the Automation Framework

ThinkSys then developed a reusable automation framework for browser, Windows desktop, and AI-driven workflows.

The automation stack included:

  • Playwright with Python for browser and workflow automation
  • Pywinauto and PyWin32 for Windows desktop automation
  • Promptfoo for AI and LLM evaluation
  • GitHub for version control
  • GitHub Actions for execution
  • Pytest reports for test reporting

Automation was focused on stable, repeatable, and regression-heavy workflows rather than trying to automate every scenario.

Step 5: Validating AI and CAD/CAE Workflows

AI-driven workflows required a different testing approach because a valid response could vary between executions.

Instead of checking only exact wording, the team evaluated whether the response met defined acceptance criteria and whether the correct action followed.

Testing covered:

  • Prompt and response behavior
  • Expected outcomes
  • Acceptable response variations
  • Negative scenarios
  • Model comparisons
  • Agent actions
  • Final application results

The end-to-end flow was validated as:

Prompt → AI Response → Agent Action → External Application State → Expected Result

This ensured that testing did not stop when the model produced a reasonable response. The final engineering result also had to be correct.

Step 6: Structuring Defects and Running Regression Cycles

ThinkSys introduced a structured defect-management process so both teams could work from the same view of product quality.

Each defect received a unique ID and was documented with the relevant details, including severity, status, and supporting information.

Issues were grouped by severity so business-critical defects could be addressed before lower-impact problems.

A shared dashboard gave Cosmon visibility into the complete defect lifecycle, including:

  • Open defects
  • Severity
  • Fix status
  • Retesting
  • Closed issues
  • Overall QA progress

Regression cycles were then used to verify fixes and check whether changes to the application, model, prompt, agent behavior, or integration introduced new issues.

Conclusion

With targeted efforts and a strong QA framework, Cosmon witnessed 90% more stable releases in their engineering applications. If you want the same level of confidence in your product, ThinkSys can help you build a QA process focused on finding critical issues before your users do.

With our Zero Critical Bug Guarantee, we help identify, prioritize, fix, and validate business-critical defects before release. Get better visibility into product quality, reduce release risks, and ship with confidence.

If you want similar results for your application, schedule a call with ThinkSys today.

Facing similar challenges to Cosmon?
 

Gaurav Mehta

About the Author

Gaurav Mehta

Experienced Certified Scrum Master and QA Lead with 12+ years of expertise in Agile delivery, software quality assurance, team leadership, and stakeholder management. Guiding cross-functional Scrum teams through planning, execution, and continuous improvement while ensuring the delivery of high-quality software solutions. Passionate about fostering Agile best practices and leveraging Artificial Intelligence in software testing to optimize processes, enhance productivity, and improve software quality.

Table of Contents

  • Meet Cosmon
  • The Challenge
  • The Proposed Solution
  • Results
  • Implementation Steps
  • Conclusion