
05.02.2024
In Part 2 of the „Testing AI Systems“ series, Tom looks at the quality characteristics that are particularly important when evaluating AI systems. He uses the relevant ISTQB certification as a guide.
I'm previous chapter I have described my first steps in the field of AI testing – from image classification to my own AI-assisted translator. Now we dive deeper into the world of testing AI systems and take a look at the characteristics of the systems to characterise their suitability as translators.
The ISTQB „Certified Tester Foundation Level“ has established itself as an international standard and offers a solid foundation for software testing. The specialist module „Certified Tester AI Testing“ examines the quality characteristics that are particularly important when evaluating AI systems. These characteristics help to assess the performance, reliability, and efficiency of AI systems.
After familiarising myself with the AI-specific features, my aim was to apply them to the translators „Libre Translate“ and „DeepL“to apply and to compare both systems with each other. Both AI translators have already been in previous chapter presented in terms of content.
When comparing Libre Translate and DeepL, Flexibility and adaptability crucial factors. Libre Translate proves flexible in the area of text-based translation and speech recognition. It is based on an OpenNMT model trained with Argos Translate and allows for the training of additional languages with tools such as Locomotive. In contrast, DeepL offers flexibility for different types of texts and has advanced features for eloquent text revisionDeepL WriteI cannot assess the system's adaptation to a new context, as DeepL has a black-box nature, meaning it's not clear exactly how the system works.
Both systems can text autonomous and understand the context, with DeepL exhibiting better contextual understanding. The translators also offer APIs for integration into various systems. Nevertheless, there is a possibility of incorrect translations, especially with specialised terms.
Regarding the Evolution Do both systems differ. Libre Translate is not a self-learning system and requires training for new languages and contexts. In contrast, DeepL can improve its translation capabilities over time, but also relies on training.
With regard to Distortions Both systems show a so-called gender bias, with DeepL performing better in the test. Details on this will follow in the next post. Gender bias in translations often manifests in the use of gender-specific terms and formulations, which can reinforce traditional role stereotypes. When translating various swear words from German into English and vice versa, I could not identify any systematic limitations in translating specific terms.
Australia ethically Libre Translate promotes sustainable development and helps overcome language barriers. The system is transparent and open-source. DeepL has similar ethical principles, although it is less transparent as a black-box system.
Both systems show Side effects how gender bias, whereby DeepL allows for alternative translations and fine-tuning of the output texts. Both systems are immune to Reward hacking. A possible danger here nonetheless would be that the systems do not reflect the full interaction in terms of content, i.e. do not provide a 1:1 translation, but nevertheless convey the essential facts.
Regarding Transparency, interpretability, and explainability offers LibreTranslate, through its open-source nature, the possibility to look at and understand the system more precisely. DeepL, however, as a closed-source system, strives for transparency through API documentation and Blog article, which provide insights into how the system works.
The Functional safety is guaranteed with both systems. They are robust and adhere to security protocols during data transmission (HTTPS encryption). In comparison, Libre Translate does not always provide reliable translations, as could already be read in the first post: “Summer holiday in Turkey”. In safety-critical applications, Libre Translate would therefore not be a suitable choice.
Generally, when evaluating AI systems, it is important to consider all of the quality characteristics mentioned above. Each user has different preferences and places different emphasis on these characteristics.
While AI translators are being compared, Libre Translate scores points with its transparent, open-source structure. DeepL, on the other hand, offers advanced features and self-learning capabilities, albeit with less transparency. At first glance, both systems equally meet the criteria of ISTQB-AI quality characteristics. But how can it be more precisely verified that the systems do not deliver faulty translations?
When testing software, it is important to define clear test objectives in order to systematically search for potential errors in the software. This also applies to systems with AI. Specifically for AI translators, I have set the following test objectives for further evaluation:
The presented test objectives provide a solid framework for further evaluation of AI translators. In the next chapter I will provide a detailed insight into the creation process of concrete test cases and how I implemented and executed them in Open Text ALM Octane. This outlook will therefore offer a deeper insight into the practical application of „ISTQB AI“ in the context of AI testing.

Arrange an initial consultation