10.00 8.30 8.29 English German
10.00 8.30 8.29 English German

Artificial Intelligence (AI)

This article describes the use of artificial intelligence (AI) in the Vienna Test System, provides information on potential risks associated with AI in the context of the Vienna Test System, and offers recommendations on how users can mitigate these risks.

Use of AI in the Vienna Test System

Currently, no artificial intelligence is used in the Vienna Test System. This applies to the content, presentation and scoring of the tests, the further processing and export by the Vienna Test System (e.g. in the form of automatically generated results reports), as well as the administration and processing of person and usage data. The VTS is therefore not an AI system as specified by the EU AI Act and is not subject to the specific requirements of the AI Act.

Irrespective of this, users of the VTS are responsible for complying with the applicable legal, professional ethical and technical requirements for psychological assessment in order to assess whether, and in what form, the use of the VTS is permissible.

Please also see Certifications and reference models

Use of external AI systems

Particular caution is required when manually using external AI systems such as ChatGPT or Claude in conjunction with the Vienna Test System (VTS). Even though the VTS itself does not use AI, the further processing of test content, test results or personal information in external AI systems may give rise to legal and professional risks. Users must therefore check in advance whether the intended use is compatible with the applicable legal, data protection, professional and ethical requirements. This also applies to browser-integrated AI, translation or assistance functions, unless it is clearly established that the processing is carried out confidentially and in compliance with data protection regulations.

Personal data, test results, diagnostic information or other confidential details relating to test takers should, as a general rule, not be entered into AI systems (see GDPR or the relevant national data protection guidelines). This includes, in particular, names, dates of birth, diagnoses, psychological findings, personality traits, the content of expert reports or other information that directly or indirectly identifies individuals. Even anonymised individual case data may, when combined with other information, allow conclusions to be drawn about specific persons and should therefore be handled with appropriate caution.

Similarly, protected and sensitive test materials must not be entered into external AI systems. These include, in particular, test items and stimulus materials, test instructions, scoring schemes, norm tables, manual contents, as well as screenshots from test procedures and other copyright-protected content. The input of such materials into AI systems is prohibited, as it jeopardises the security of psychological tests and may constitute a copyright infringement. This also applies even if the input is solely for the purpose of asking questions, summarising texts or obtaining assistance with interpretation.

The use of external AI systems for scoring, interpretation or evaluation of test results, as well as for the preparation of psychological reports or expert opinions, is not recommended. Decisions, professional assessments and conclusions must be made personally by suitably qualified psychologists, in a scientifically sound and context-sensitive manner. AI-generated content may be inaccurate, incomplete or unsuitable and must not replace professional responsibility. Use for linguistic, structural or formal support may be considered only if no personal, confidential or protected test content is processed and all content is subject to full professional review.

Use of AI by test takers

The widespread availability of AI applications poses a challenge for digital psychological assessment, as test takers may use AI to prepare for, receive assistance with or edit tests and questionnaires independently. This can limit the validity of psychological test results, particularly because AI-generated responses resemble human responses and cannot be reliably detected either in terms of content or through data analysis (Panizza et al., 2026; Westwood, 2025; Walker et al., 2026).

Text-based, easily copyable or linguistically formulated tasks—such as open-ended responses, questionnaires, knowledge tests and multiple-choice questions—are particularly affected. The risk is lower for complex visual, spatial, time-critical or interactive tasks (Abdelkarim et al., 2025). Studies show that AI is used in open-ended response formats, though this is not necessarily linked to a deliberate intention to cheat (S. Zhang et al., 2025), and that knowledge tests tend to benefit more from AI than tests of fluid intelligence (Stelling et al., 2026).

Although technical options for detecting and preventing the use of AI are under discussion (Martherus et al.; 2025; Asher et al., 2026; D. Zhang et al., 2025), no corresponding measures have currently been implemented in the VTS. The risk therefore depends largely on the chosen test mode: it is highest in the unsupervised open mode, lower in the online supervised proctored mode, and lowest in the on-site supervised controlled mode, provided that the conditions are consistently monitored (see Testing modes).

What can I do to ensure a high test security?

Please refer to the page Ensuring a high test security to find a summary of possible measures that can be used to improve test security.

References

Abdelkarim, S., Lu, D., Flores, D.-L., Jaeggi, S., & Baldi, P. (2025). Evaluating the intelligence of large language models: A comparative study using verbal and visual IQ tests. Computers in Human Behavior: Artificial Humans, 5, Article 100170. https://doi.org/10.1016/j.chbah.2025.100170

Asher, M. W., Gold, G., Chen, E., & Carvalho, P. F. (2026). Chatbots are undermining crowdsourced research in the behavioral sciences: Detecting artificial intelligence–assisted cheating with a keystroke-based tool. Advances in Methods and Practices in Psychological Science. https://doi.org/10.1177/25152459261424723

Martherus, J., Podkul, A., Cook, E., & Liebowitz, R. (2025). How to detect AI-assisted interviews in online surveys. Survey Practice, 18. https://doi.org/10.29115/SP-2025-0016

Panizza, F., Kyrychenko, Y., & Roozenbeek, J. (2026). Survey-taking AI tools surpass human abilities. Here’s what we can do about it. Nature, 650(8101), 293–295. https://doi.org/10.1038/d41586-026-00386-2

Stelling, H., Kraus, A., Grieb, G., & Güler, I. (2026). Performance of large language models on cognitive aptitude testing: A multi-run evaluation on the German Medical School Admission Test (TMS). European Journal of Investigation in Health, Psychology and Education, 16(2), Article 23. https://doi.org/10.3390/ejihpe16020023

Walker, D. C., Tran, M. P. N., Bizer, G. Y., & Flynn, S. T. (2026). Initial attempts to detect or screen out AI responses prove elusive in the age of agentic AI. International Journal of Eating Disorders, 59(4), 666–673. https://doi.org/10.1002/eat.70024

Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. Proceedings of the National Academy of Sciences, 122(47), Article e2518075122. https://doi.org/10.1073/pnas.2518075122

Zhang, S., Xu, J., & Alvero, A. J. (2025). Generative AI meets open-ended survey responses: Research participant use of AI and homogenization. Sociological Methods & Research, 54(3), 1197–1242. https://doi.org/10.1177/00491241251327130

Zhang, D., Katoh, M., & Pei, W. (2025). Detecting the use of generative AI in crowdsourced surveys: Implications for data integrity. arXiv. https://doi.org/10.48550/arXiv.2510.24594