The civil aviation industry presents a fascinating paradox. On one hand, it operates under some of the world’s most stringent safety requirements and highest barriers to entry, having successfully evolved into one of the safest transportation and engineering systems in human history. On the other hand, it is also one of the earliest sectors where automation technology achieved monumental success—from autopilots decades ago to automated instrument landings, pilots and air traffic controllers have long been accustomed to handing control over to deterministic algorithms. However, as modern AI—built on probabilistic models and black-box architectures—attempts to enter this domain, the overarching question remains: Can AI replicate the legendary success of classic flight automation?
According to a TMTPost 2025 survey, as many as 60% of AI proof-of-concept (PoC) projects in civil aviation remain stuck in laboratories, struggling to make it into actual production environments. Many technical teams and industry professionals are often puzzled: with compute power and model parameter counts growing exponentially, why is it still so difficult to send an algorithmic command directly to a flight control computer at 30,000 feet, or replace a tower controller in issuing a verbal clearance? The root cause has never been that models aren’t smart enough, but rather that civil aviation imposes exceptionally strict requirements on airworthiness certification determinism and scenario risk classification.
To guarantee safety at 30,000 feet, civil aviation software must strictly comply with the DO-178C airworthiness standard. Its highest safety level, DAL A, mandates that the catastrophic failure rate of a system must be below per flight hour. Every single line of traditional flight control code undergoes rigorous formal verification and deterministic testing; given specific inputs, it invariably produces a unique, predictable output. However, artificial intelligence models grounded in probability and statistics are fundamentally black boxes. The inherent randomness of their outputs collides directly with the deterministic guarantees required for airworthiness certification. On the physical isolation front, the DO-326A standard further enforces strict separation between the Aircraft Control Domain—which directly impacts flight safety—and the Information Services Domain using one-way security gateways.
The case of the airborne collision avoidance system ACAS Xu vividly illustrates this conflict. MIT Lincoln Laboratory once attempted to compress massive collision avoidance lookup tables by 1,000 times using neural networks to save onboard computational resources. However, when researchers tested safety boundaries using formal verification tools like Reluplex (developed by Stanford) and NNV, the published study revealed that under specific two-aircraft convergence angles, the neural network would suddenly issue erroneous pitch-down turn commands, directly leading to mid-air collisions. This finding sounded an alarm for the industry: neural networks without formal proof simply cannot directly take over DAL A core flight control systems.
A similar dilemma unfolded in air traffic control (ATC) towers. In the European SESAR MALORCA project, researchers tried using machine learning to recognize air-ground voice communications between controllers and pilots to assist controllers in issuing instructions. According to the DLR/Idiap academic study, the project team reduced command recognition error rates in real operational environments at Prague and Vienna airports to between 0.9% and 4.0% (0.9% for Prague, 4.0% for Vienna), and 0.5% to 2.5% in laboratory settings. However, even a 0.9% misinterpretation rate in busy airway management could potentially trigger a major air disaster. Consequently, the MALORCA system was ultimately restricted to a Level 1 advisory role, serving merely as visual prompts for controllers to manually review and confirm via keypresses, with no algorithm allowed to automatically issue voice or datalink commands to aircraft.
Faced with airworthiness challenges stemming from probabilistic black boxes, EASA chose the most rigorous and systematic path: constructing airworthiness certainty through hard regulations and standard frameworks. EASA published the world’s first official document to integrate AI airworthiness certification into technical guidance, Concept Paper Issue 2. This 283-page publication divides civil aviation AI operational authority into a six-level framework ranging from Level 1A to Level 3B, providing a clear roadmap for AI deployment in civil aviation systems: - Level 1A/1B Assistance Mode: AI acts as an advisor, primarily responsible for information acquisition, analysis, and decision support, while humans perform all operations and retain final accountability. - Level 2A/2B Collaboration Mode: AI becomes a human partner, executing tasks automatically upon human instruction, or automatically selecting and executing actions under real-time human supervision. - Level 3A/3B Autonomous Mode: The system achieves autonomous decision-making and execution backed by warning safeguards, eventually advancing toward unsupervised, full autonomy.
To give these guidelines legal force, EASA further issued the NPA 2025-07 notice, aligning civil aviation requirements with the European Union’s broader AI framework, the EU AI Act (Regulation (EU) 2024/1689). Meanwhile, the ARP6983 and ED-324 Issue 1 airworthiness standards, jointly developed by EASA, SAE, and EUROCAE, are slated for official release in June 2026. This standard serves as an airworthiness engineering manual custom-tailored for machine learning, establishing clear technical lifecycle requirements for training, testing, and verifying non-linear models. According to EASA’s roadmap, the world’s first airworthiness approval for a machine learning component based on high-complexity computer vision (Design Assurance Level DD) is expected in early 2026, while certification for fully autonomous Level 2 and Level 3A operations is planned around 2035. Europe’s seemingly intricate regulatory process actually supplies the entire aviation manufacturing industry chain with invaluable R&D expectations and deterministic baselines.
In contrast to Europe’s comprehensive rulemaking strategy, the FAA adopted a more moderate set of guidelines. In its Roadmap for AI Safety Assurance Version I, the FAA outlined seven core principles and clearly distinguished between safety of the AI system itself and AI for safety enhancement. Instead of rushing to enact mandatory new regulations, the FAA chose to wait for industry standards to mature organically. However, this wait-and-see regulatory stance has also had drawbacks: as highlighted in the FAA REDAC 2023 report, the lack of clear airworthiness certification certainty has left the US aviation industry largely hesitant to integrate artificial intelligence into critical flight control products.
However, in domains far removed from core flight controls—such as maintenance, repair, and overhaul (MRO), flight operations dispatch, and scheduling—the US market has demonstrated striking commercial agility by establishing a clear financial ledger for return on investment (ROI). General Electric’s Predix platform monitors engine health for more than 4,600 commercial aircraft worldwide, while Airbus’s Skywise platform connects real-time operational data from over 5,000 aircraft globally, preventing numerous unscheduled maintenance flight cancellations through predictive warnings. On the flight dispatch and scheduling side, United Airlines’ ConnectionSaver system has facilitated over 3.3 million successful passenger connections since its 2019 launch, saving 54,000 connections in 2025 alone. Alaska Airlines’ Flyways route optimization system saved 480,000 gallons of fuel for a single carrier, while Qantas’s Constellation optimization platform generates $92 million in cost savings annually.
In airport surface safety, Honeywell’s SURF-A and SmartX runway incursion prevention systems have been successfully installed and deployed across Southwest Airlines’ fleet of more than 700 Boeing 737 aircraft. These successful implementations are concentrated in areas with controllable risks, such as maintenance operations and decision support. The US market has proven through action that as long as the economic equation works in low-risk scenarios, AI can rapidly achieve commercial viability at scale.
Civil aviation in China has demonstrated distinct policy-driven dynamics and scenario-specific breakthroughs in AI adoption. In November 2025, the CAAC officially issued the Implementation Opinions on Promoting the High-Quality Development of “Artificial Intelligence + Civil Aviation”, outlining 42 typical application scenarios spanning flight safety, maintenance, and airport operations. Aligned with phased targets for 2027, 2030, and 2035 established in the previously released 14th Five-Year Roadmap for Smart Civil Aviation Development, CAAC is leveraging robust industrial policy to drive deep integration between data governance standards like MH/T 5057-2021 and practical scenario deployments.
This policy guidance has yielded notable results in airport operational perception and security screening. The AI-powered security inspection system developed by Nuctech has been deployed at Chengdu Tianfu International Airport and several major global hubs, boosting overall security lane throughput efficiency by 12.2%. At Hong Kong International Airport, the DATMS digital tower and AFODDS runway debris and incursion warning system—deployed in partnership between Searidge and NATS—achieved high-precision, real-time sensing of foreign object debris (FOD) and runway incursion risks. According to a March 2026 industry report by Claims Journal, such perception systems have substantially enhanced major airports’ operational resilience under adverse weather conditions.
Concurrently, airports and airlines in regions like Shenzhen and Xinjiang have begun actively exploring localized operational deployments of frontier Large Language Models (LLMs). By integrating domestic models such as DeepSeek or Qianrang, operations teams have constructed offline safety warning systems and intelligent Q&A knowledge bases centered around Flight Crew Operating Manuals (FCOM) and Minimum Equipment Lists (MEL). Although these initiatives remain in an offline advisory stage, they have generated valuable field experience for Chinese civil aviation in rapid technical document retrieval and operational decision support.
Reflecting on the practical explorations across Europe, the US, and China, a production-grade paradigm for civil aviation AI deployment comes into clear focus. Taking the European Organisation for the Safety of Air Navigation as an example, the traffic prediction AI developed by EUROCONTROL successfully reduced air traffic flow prediction errors by 71% in production environments, automatically processing over 30,000 flight plans daily and elevating automated processing rates from 96.3% to 99.5%. Additionally, the ALIX baggage image-matching system, jointly developed by SITA and IDEMIA, has been seamlessly integrated across multiple international airports. These successfully deployed production systems all adhere to a shared engineering philosophy: Shadow Mode execution and Human-in-the-Loop architecture.
Under Shadow Mode, algorithms ingest real-time sensor and system data in the background to calculate optimal recommendations, yet issue zero control commands until proven over millions of fault-free operational hours. The authority to confirm any critical decision remains strictly with human controllers, captains, or flight dispatchers. By providing incremental assistance without usurping core control authority, this design gives algorithms practical experience in live operational settings while preserving an unyielding safety baseline for civil aviation.
However, Human-in-the-Loop mechanisms introduce a distinct psychological challenge: automation bias. An academic study published in Springer indicates that when controllers rely on highly accurate automated tools over extended periods, should the automated system occasionally suffer a missed detection, the manual hazard detection rate without system prompts was actually 15.56% higher compared to a purely manual baseline (Source: MDPI Aerospace 2021, https://www.mdpi.com/2226-4310/8/9/260). This implies that future civil aviation AI interface design cannot merely display recommended answers; it must incorporate interaction mechanisms specifically engineered to maintain human operators’ situational awareness, preventing blind trust and cognitive complacency.
For the foreseeable future, intelligence in civil aviation will not mean replacing human pilots or controllers overnight with an algorithm; rather, it represents a long-term engineering compromise. Europe’s top-level design in airworthiness standards, America’s ROI validation in commercial scenarios, and China’s rapid pilot deployments in operational perception together compose a complete landscape of civil aviation AI development. Only by respecting the deterministic boundaries of airworthiness certification and honoring the physical and psychological constraints of safety can artificial intelligence truly transition from laboratories into towers and onto runways, striking firm roots in the sky.