We can now generate software faster than we can believe it.
Generative AI is not autonomy. It does not give a system authority to decide or act. But its arrival in the safety-critical build chain is the most consequential recent change in how we build and scale aerospace autonomy. It has changed the rate at which we can produce code, tests, documentation, and designs that enter the safety-critical build chain. Production can now move faster than the processes we use to justify confidence in what we build.
But development is only one clock in the larger process. Scaling autonomy requires justified confidence to keep pace at three different timescales: while we develop the system, while it decides and acts, and while we learn from operations.
Before Deployment: The Development Clock
To see the development problem clearly, we first have to stop treating every new technology as a synonym for autonomy. Automation executes a defined task. An autonomous system has delegated authority to choose and act toward a goal within defined bounds. Autonomy is a system-level property, not a particular technology inside the system.
Artificial intelligence is a broad family of theories and methods for perception, reasoning, learning, and decision support. Machine learning is one part of that family: it creates models that fit behavior from data rather than specifying every rule directly. Learned models can perform functions for which explicit rules are impractical, but their behavior is less tractable to conventional assurance. Both learned and rule-based methods may support an autonomous decision, but neither is autonomy. Some autonomous systems contain no machine learning. Many systems use machine learning but have no authority to act.
Human-machine teaming brokers how authority, information, and responsibility pass between people and systems. Digital engineering, digital twins, and the digital thread address how we build systems, organize models, and connect evidence across the lifecycle. They can support autonomy and its assurance, but they do not make a system autonomous.
All these terms can get conflated because they appear to put a machine where people expect a human. However, the distinctions matter because the technologies create different assurance problems. Generative AI increases the volume of engineering artifacts that people must review and defend. Learned onboard models present a separate problem: they challenge the form of evidence itself.
Aerospace has traditionally kept roughly two assurance ledgers. Physical components can fail randomly during operation as they encounter damage, degradation, and environmental stresses. We model those failures over operating time and allocate their rates against system safety targets. Software does not wear out; its failures arise from design errors and recur when the triggering inputs and state recur. So we control the process that produces it through requirements, traceability, coverage, testing, and review.
A learned model fits neither ledger comfortably. It is engineered like software, but much of its behavior is fitted from data rather than traced from written requirements. Its failures are usually systematic over parts of the input space rather than random in time, so a time-based failure rate does not tell us where the model will break. Yet conventional process evidence cannot trace billions of learned parameters to individual requirements. The problem is not simply more software to verify. Our established assurance machinery does not fully describe the artifact we need to defend.
At the Point of Action: The Decision Clock
One practical bridge is to put the uncertain component “in a box.” An independently assured runtime monitor checks its output before it can affect the vehicle, with a safe fallback when the check fails. The monitor does not prove that the learned model is correct. It supports a narrower claim: under stated assumptions, the system can remain within specified safety bounds.
But that claim has a deadline. A planning decision with seconds available may permit a richer check than a perception decision due in milliseconds. Runtime assurance is therefore a time and computation budget: how much can we check before the system must act?
The budget becomes more consequential as autonomy scales. Scaling can mean more delegated authority per vehicle, more vehicles per human supervisor, or moving autonomous decisions from an onboard function into fleet operations and airspace management. Each dimension increases the number, variety, or pace of decisions that demand confidence. The right level of rigor must follow the consequence and the operating domain, rather than the presence of an “AI” label.
After Deployment: The Learning Clock
No predeployment process will discover every failure, near miss, or exit from an operational design domain. Fielded systems produce evidence that tests and simulations miss. Yet each organization sees only a fraction of those cases. If everyone learns alone, the industry will repeat mistakes and one company’s failure can slow deployment for everyone.
Aviation has faced this problem before. So it built protected reporting, shared operational data, neutral analysis, and industry-government teams that turn observations into fixes without asking competitors to give away the capabilities on which they compete. The working question for autonomous operations is what similar machinery would require: a common event taxonomy, a protected data commons, clear data boundaries, and a neutral analyst that returns useful lessons to participants.
This is where AIAA’s Scaling Autonomy and Autonomous Operations Working Group comes in. AIAA technical committees have addressed autonomy for years, with the Intelligent Systems Technical Committee playing a sustained leadership role. AIAA’s 2025 autonomy paper deliberately focused on single-vehicle, onboard autonomy and left fleet and operational scale outside its scope. The 2025 CTO Summit created our group to pick up there.
We have assembled operators, technology leaders, regulators, and researchers from aviation, uncrewed aircraft systems, and space. Our three-year horizon gives us room to move beyond another survey and develop concrete asks for AIAA, standards bodies, regulators, and industry. A 2027 white paper will be one step in that work.
The hard question is not whether organizations should share everything. They should not. It is what protections and value would make them willing to share the failures from which the entire field needs to learn.
Compete on capability; share the failure modes.

