Professional IT services from accompio for companies in Germany.
Blog

Testing AI: focus on quality features

05.02.2024

In Part 2 of the „Testing AI Systems“ series, Tom looks at the quality characteristics that are particularly important when evaluating AI systems. He uses the relevant ISTQB certification as a guide.

Two men discuss quality characteristics when testing AI systems.

I'm previous chapter I have described my first steps in the field of AI testing – from image classification to my own AI-assisted translator. Now we dive deeper into the world of testing AI systems and take a look at the characteristics of the systems to characterise their suitability as translators.

On the trail of ISTQB quality characteristics

The ISTQB „Certified Tester Foundation Level“ has established itself as an international standard and offers a solid foundation for software testing. The specialist module „Certified Tester AI Testing“ examines the quality characteristics that are particularly important when evaluating AI systems. These characteristics help to assess the performance, reliability, and efficiency of AI systems.

  • Flexibility and adaptabilityThe system should be able to adapt easily to different situations and environments, including new and unforeseen ones.
  • Autonomy in AI SystemsThis is about a system being able to operate autonomously for a longer period of time. The creator of the system must define how long and under what conditions.
  • EvolutionA system should be able to improve itself when its environment changes. This is particularly important for self-learning AI systems.
  • DistortionSometimes, the results of AI systems can deviate from what is considered fair. It is important to ensure that such deviations are controlled, for example, with regard to gender or income.
  • Ethics in AI systemsAI systems should follow clear rules that ensure they serve humanity, respect democratic values, and are transparent.
  • Side effects and reward hackingIf things are overlooked during development, undesirable effects can occur. An example of this is a translator who mistranslates legal terms, leading to legal misunderstandings. Reward hacking is the attempt by a system to achieve goals in an intelligent way that may contradict the original intentions of the developers.
  • Transparency, interpretability, and explainabilityThis is about how easily users can understand how the system works and why it produces certain results.
  • Functional Safety and AIIt must be ensured that AI systems reliably fulfil their tasks without unexpected errors, especially in safety-critical applications.
The ISTQB syllabus for „AI Testing“ includes an entire chapter on quality characteristics.

Quality Features Compared: A Look at LibreTranslate and DeepL

After familiarising myself with the AI-specific features, my aim was to apply them to the translators „Libre Translate“ and „DeepL“to apply and to compare both systems with each other. Both AI translators have already been in previous chapter presented in terms of content.

When comparing Libre Translate and DeepL, Flexibility and adaptability crucial factors. Libre Translate proves flexible in the area of text-based translation and speech recognition. It is based on an OpenNMT model trained with Argos Translate and allows for the training of additional languages with tools such as Locomotive. In contrast, DeepL offers flexibility for different types of texts and has advanced features for eloquent text revisionDeepL WriteI cannot assess the system's adaptation to a new context, as DeepL has a black-box nature, meaning it's not clear exactly how the system works.

Both systems can text autonomous and understand the context, with DeepL exhibiting better contextual understanding. The translators also offer APIs for integration into various systems. Nevertheless, there is a possibility of incorrect translations, especially with specialised terms.

Regarding the Evolution Do both systems differ. Libre Translate is not a self-learning system and requires training for new languages and contexts. In contrast, DeepL can improve its translation capabilities over time, but also relies on training.

With regard to Distortions Both systems show a so-called gender bias, with DeepL performing better in the test. Details on this will follow in the next post. Gender bias in translations often manifests in the use of gender-specific terms and formulations, which can reinforce traditional role stereotypes. When translating various swear words from German into English and vice versa, I could not identify any systematic limitations in translating specific terms.

DeepL offers alternative word suggestions by clicking on individual words, even gender-specific ones here.


Australia ethically Libre Translate promotes sustainable development and helps overcome language barriers. The system is transparent and open-source. DeepL has similar ethical principles, although it is less transparent as a black-box system.

Both systems show Side effects how gender bias, whereby DeepL allows for alternative translations and fine-tuning of the output texts. Both systems are immune to Reward hacking. A possible danger here nonetheless would be that the systems do not reflect the full interaction in terms of content, i.e. do not provide a 1:1 translation, but nevertheless convey the essential facts.

Regarding Transparency, interpretability, and explainability offers LibreTranslate, through its open-source nature, the possibility to look at and understand the system more precisely. DeepL, however, as a closed-source system, strives for transparency through API documentation and Blog article, which provide insights into how the system works.

The Functional safety is guaranteed with both systems. They are robust and adhere to security protocols during data transmission (HTTPS encryption). In comparison, Libre Translate does not always provide reliable translations, as could already be read in the first post: “Summer holiday in Turkey”. In safety-critical applications, Libre Translate would therefore not be a suitable choice.

Generally, when evaluating AI systems, it is important to consider all of the quality characteristics mentioned above. Each user has different preferences and places different emphasis on these characteristics.

While AI translators are being compared, Libre Translate scores points with its transparent, open-source structure. DeepL, on the other hand, offers advanced features and self-learning capabilities, albeit with less transparency. At first glance, both systems equally meet the criteria of ISTQB-AI quality characteristics. But how can it be more precisely verified that the systems do not deliver faulty translations?

Key aspects in quality assessment

When testing software, it is important to define clear test objectives in order to systematically search for potential errors in the software. This also applies to systems with AI. Specifically for AI translators, I have set the following test objectives for further evaluation:

  • Completeness checkEnsure that the translation captures the entire content of the original text, without omitting any crucial information.
  • Comprehensibility checkEnsure the translation is easily understandable, so users can grasp the content in the target language without difficulty.
  • Accuracy checkEnsuring that AI translation is precise and accurate, to avoid misunderstandings or misinterpretations.
  • Dealing with technical termsEnsure that the system correctly recognises specialist terminology and uses it appropriately in the translation.

Outlook: Test cases in ALM Octane

The presented test objectives provide a solid framework for further evaluation of AI translators. In the next chapter I will provide a detailed insight into the creation process of concrete test cases and how I implemented and executed them in Open Text ALM Octane. This outlook will therefore offer a deeper insight into the practical application of „ISTQB AI“ in the context of AI testing.

Woman with a headset in customer service at Accompio IT Services.

Get in touch with us

We at accompio will be happy to help you.

Arrange an initial consultation

This field is for validation purposes and should be left unchanged.
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form

From time to time we would like to inform you about our products and services as well as other content that may be of interest to you. You can unsubscribe from these communications at any time. If you agree to us contacting you for this purpose, please tick the following box. You can revoke your consent at any time with effect for the future - via the unsubscribe link at the end of each e-mail or by e-mail to info@accompio.com.

We process and store your data. You can find further information at Privacy Policy.