← Back to feed
2026-07-21alignmentinfra

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Gjergji Kasneci, Enkelejda Kasneci

PDF preview for The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
Read on arXiv →

Key claim

A five-layer framework for diagnosing AI safety risks.

In plain English

Imagine you're building an AI system that interacts with users and makes decisions based on their inputs. While it's easy to spot obvious failures, like a chatbot giving harmful advice, the real challenge lies in the subtle, systemic issues that can arise over time. These include problems like overreliance on certain data sources or the gradual erosion of trust in the system, which can go unnoticed until it's too late. This is what's called epistemic integrity, and it highlights how critical it is to ensure that the AI's decision-making process is transparent and accountable. Other failure modes include control integrity, where the system's permissions might be compromised, and temporal integrity, where safety measures might not hold up as the system evolves or is updated.

To tackle these hidden risks, the authors propose a five-layer framework that helps diagnose and address these issues. Each layer focuses on a different aspect of the socio-technical system surrounding AI, from how evidence is represented to the robustness of organizational oversight. By identifying under-recognized risk patterns, such as prompt injection and memory poisoning, this framework shifts the focus from merely evaluating model performance to ensuring the entire system is reliable and safe. For builders, this means that instead of just testing for obvious failures, you now have a structured way to think about the broader implications of your AI systems and how to design them for long-term safety and accountability.

Novelty
8.0/10

The framework addresses under-explored safety risks in AI systems.

Reliability
7.5/10

The proposed framework is well-structured and identifies specific risk patterns.

Deep reliability assessment

The methodology supports a framework for diagnosing hidden risks in AI systems, focusing on socio-technical reliability rather than just model-centric evaluation. It does not provide empirical prevalence data but organizes concerns into a structured framework.

Reproducibility

no

Key figure

Figure 1 illustrates the concept of AI safety as an iceberg, where visible failures are just the tip, and the more consequential challenges are submerged across five layers of socio-technical integrity.