The South Korean government is pursuing domestic technology independence in artificial intelligence (AI) safety and security, which it views as a future strategic technology, while also seeking cooperation with global big tech firms. The plan is to expand cooperation on safety evaluation and cybersecurity with global AI companies such as OpenAI and Anthropic, while at home accelerating efforts to secure the original technology needed to develop tools that prevent hallucinations and misuse in AI models.

According to industry sources Tuesday, the Institute of Information & Communications Technology Planning & Evaluation (IITP) and the Ministry of Science and ICT (MSIT) will pursue a new "AI Safety and Trust Technology Development Project" worth 300 billion won over eight years, from as early as 2027 through 2034. To this end, IITP recently issued a public notice for a "preliminary planning research service for a large-scale R&D project on AI Safety and Trust."
The project aims to evaluate the safety of AI models such as large language models (LLMs) at the national level, and to secure technology that can block dangerous responses, hallucinations, and malicious misuse. Currently, global LLMs such as ChatGPT, Claude, and Gemini have their safety verified independently by each company, but national-level testing and certification systems are not sufficient.
As a result, while risks such as the provision of harmful information, malware generation, automated phishing, and disinformation creation are growing worldwide, responding to them is not easy. Even when each model developer takes measures to block searches for harmful information, users can easily abuse AI tools if they find workaround methods that companies cannot detect, depending on their cultural context. "When users ask how to manufacture drugs using indirect expressions, some models provide dangerous information," an IITP official said. "We need a standard system that can evaluate the safety of each model."
The government is laying the foundation for its own research and development while also expanding its global cooperation network. The MSIT signed a memorandum of understanding with Anthropic on Tuesday in the areas of securing AI safety and cybersecurity. The two sides agreed to analyze the impact of AI on cyberattacks and defense, and to evaluate the safety and misuse risks of AI models in the Korean-language context. Red team evaluation of autonomous AI agents, discovery of AI vulnerabilities, and sharing of cyber threat information are also subjects of cooperation. Earlier, on Monday, the AI Safety Institute signed a memorandum of understanding with OpenAI, agreeing to share knowledge and best practices on safety evaluation methodologies and benchmarks across high-risk areas. The two sides will exchange technical information to develop an AI safety evaluation system that reflects the context of Korean society, and will also cooperate on establishing an internationally applicable evaluation system.
Experts point out that, as AI safety emerges as a national security issue, securing a national-level evaluation system and original technology is urgent. "AI safety evaluation should be viewed as public infrastructure," an expert in the field of AI safety said. "We need to develop, together, an evaluation system that reflects the Korean language and domestic social and cultural contexts, defense technology that prevents misuse, and technology that can explain and verify the basis for a model's judgments."






