The AI City: From Multi-Agent Intelligence to CityAI

Otthein Herzog , Qingrui Minyag Jiang , Max Mühlhäuser , Helge Ritter , Yue Wu , Jiangkai Zhao , Hao Zhang , Klaus-Dieter Thoben , Zhiqiang Siegfried Wu

ENGINEERING Cities ››

PDF (4881KB)
ENGINEERING Cities ›› DOI: 10.2738/ENGC.2026.0010
Research Article
The AI City: From Multi-Agent Intelligence to CityAI
Author information +
History +
PDF (4881KB)

Abstract

Cities have long been conceptualized as machines, organisms, networks, and complex adaptive systems, each reflecting a computational ambition: to calculate, simulate, sense, predict, and increasingly enable cities to reason and act. Recent advances in foundation models, generative AI, large language model agents, embodied intelligence, and multi-agent systems raise a question: can intelligence emerge as an urban-scale capability? This article develops the concept of CityAI through a historical-conceptual analysis of the transition from computational urbanism and agent-based simulation to agentic urban intelligence. Building on the proposition of the AI City, it synthesizes seven recurring visions—the Technological City, Human-Centred City, Generative Planning City, Operational and Governable City, Complexity-Efficient City, Adaptive and Iterative City, and Hybrid Intelligence City—as complementary perspectives on urban intelligence. The article argues that these visions are increasingly interdependent, yet none is sufficient alone. Its central contribution is to relocate urban intelligence from models or platforms to the interactions among heterogeneous human and artificial actors. CityAI is defined as the emergent capacity of an urban system to perceive, interpret, reason, coordinate, act, evaluate, and evolve through continuous interaction among agents, infrastructures, institutions, and environments. The article concludes with a research agenda for the transition from the AI City to CityAI.

Graphical abstract

Keywords

CityAI / AI City / Multi-agent systems / Collective urban intelligence / Ultra-agent habitat

Highlight

● Traces urban intelligence from the Computed City to the CityAI era.

● Synthesizes seven global visions shaping the contemporary AI City.

● Reframes urban intelligence as relational and emergent at the urban scale.

● Distinguishes CityAI from AI for cities and multi-agent systems.

● Sets a research agenda from multi-agent intelligence toward CityAI.

Cite this article

Download citation ▾
Otthein Herzog, Qingrui Minyag Jiang, Max Mühlhäuser, Helge Ritter, Yue Wu, Jiangkai Zhao, Hao Zhang, Klaus-Dieter Thoben, Zhiqiang Siegfried Wu. The AI City: From Multi-Agent Intelligence to CityAI. ENGINEERING Cities DOI:10.2738/ENGC.2026.0010

登录浏览全文

4963

注册一个新账户 忘记密码

1 Background: Why the AI City, and Why Now?

1.1 Back to the city: rethinking urban intelligence

The relationship between cities and intelligence is older than artificial intelligence itself. Cities have always processed information [1–3]. Streets encode rules of movement; markets aggregate dispersed preferences; institutions retain collective memory; plans project desired futures; and infrastructures respond, with varying degrees of delay, to changes in demand. Long before digital computation, urban systems already displayed distributed forms of sensing, coordination, and adaptation. What has changed in the twenty-first century is not the existence of urban intelligence, but the possibility of making parts of this intelligence computationally explicit, continuously observable, and increasingly actionable.

For several decades, the dominant language for describing this transformation was the “smart city”. The smart-city paradigm connects information and communication technologies, sensor networks, data platforms, and urban services. It created a powerful agenda for making cities more measurable and more efficiently managed [4–6]. Traffic can be monitored in real time; electricity demand can be forecast; environmental conditions can be mapped; and administrative processes can be digitized. This was a decisive transition [7–9]. The city moved from periodic observation to continuous sensing. However, the smart-city paradigm is largely built around a specific relationship between technology and governance: humans set objectives, digital systems collect data, analytical models produce information, and institutions make decisions [10]. Intelligence is predominantly instrumental and task-specific.

Artificial intelligence has begun to destabilize this arrangement. Machine Learning first expanded the predictive capacity of urban analytics [11–13]. Deep Learning enabled computers to identify patterns in imagery, mobility traces, environmental data, and complex high-dimensional urban datasets. More recently, foundation models and generative artificial intelligence have introduced capabilities in semantic understanding, multimodal interpretation, content generation, and general-purpose problem solving [14,15]. Large Language Models can be embedded in agents equipped with memory, tools, goals, and action interfaces. Agentic AI systems can decompose tasks, select tools, communicate with other agents, revise plans, and act in digital or cyber-physical environments. In parallel, robotics, autonomous vehicles, Digital Twins, edge intelligence, and ubiquitous sensing are increasingly connecting computational reasoning with physical urban processes.

The resulting change is conceptual as much as technological. Earlier urban AI systems were often designed to answer questions such as: what will traffic demand be in thirty minutes? Which neighbourhood is at higher heat risk? Where should a facility be located? Which image contains a damaged road? Although these remain important questions, contemporary systems are increasingly being asked a different class of questions: What is happening across multiple urban subsystems? Why are conflicting conditions emerging? Which objectives should be prioritized? Which actors need to coordinate? What intervention is feasible under current constraints? How should a strategy be revised when residents' responses differ from planners’ expectations?

The shift can be summarized as a movement from prediction to agency. Prediction estimates future urban states [14,16,17]. Agency connects perception to objectives, reasoning, evaluation, learning and action. In urban contexts, this difference is fundamental because cities are not passive environments [18]. Every intervention changes the behaviour of people, institutions, markets, and infrastructures; these responses alter the conditions under which the next decision must be made. A traffic recommendation may redistribute congestion. A cooling intervention may reduce the urban heat island effect. A flood warning may generate evacuation flows that create new bottlenecks. A planning policy may change land values and thereby reshape development incentives. Urban actions continuously affect one another through feedback.

This problem has been approached through several emerging framings of integrated urban intelligence, including city brains, urban operating systems, Digital Twins, and the AI City. Within this broader landscape, the AI City framework proposed that artificial intelligence should not simply be added to the city as another technological layer [19]. The city requires an intelligent structure capable of perception, judgement, coordination, and response. The framework of an urban AI pivot, incorporating an urban brain, cerebellum, and central nervous system, represented an attempt to move beyond isolated applications toward an integrated urban intelligence architecture. The AI City proposition also emphasized an open ecosystem for the experimental deployment and evolution of next-generation AI technologies in urban environments. Its importance lies not only in a particular technical architecture, but in a broader conceptual move: from using AI in individual urban sectors to asking what an AI-era city itself should become.

This question has become more urgent. Urban Digital Twins are moving from static three-dimensional representations toward dynamic environments that integrate real-time sensing, simulation, prediction, evaluation, planning, and action [20–22]. Urban foundation models seek transferable representations across multiple urban data modalities and tasks. Generative urban planning systems produce spatial alternatives [23–25]. Reasoning agents are being proposed to verify rules, evaluate trade-offs, and support planning decisions. Multi-agent systems are used to represent stakeholders, simulate social behaviour, and coordinate urban services. Intelligent urban agents are becoming semi-embodied in the hybrid cyber-physical-social space of cities. At the same time, international organizations and city governments are confronting the governance implications of AI in mobility, public safety, utilities, healthcare, planning, regulation, education, and public administration.

Yet the field remains conceptually fragmented. “AI City”, “smart city”, “urban AI”, “urban intelligence”, “city brain”, “urban digital twin”, “urban foundation model”, “Agentic Urban AI”, and “intelligent urban agent” are often used in overlapping but non-equivalent ways. Some research defines progress through technical capability. Other work evaluates AI through human well-being and inclusion. Planning researchers focus on design and decision-making. Urban managers emphasize operational coordination. Economists and administrators focus on productivity and cost. Complexity researchers emphasize feedback and adaptation. Human–AI interaction researchers focus on hybrid intelligence and co-evolution. These perspectives frequently share technologies but disagree, often implicitly, about where urban intelligence resides and how it should be used.

This article proposes that the current global discourse can be productively interpreted through the seven recurring visions of the AI City. They are the Technological City, the Human-Centred City, the Generative Planning City, the Operational and Governable City, the Complexity-Efficient City, the Adaptive and Iterative City, and the Hybrid Intelligence City. The term “vision” is deliberate. The article does not claim that a systematic bibliometric procedure proves that the world contains exactly seven schools. These are not closed schools with formal doctrines or fixed memberships, and many scholars contribute to more than one. Rather, the seven-part typology is an interpretive synthesis of recurrent positions in urban AI research and practice. Each vision represents a distinct answer to a foundational question: where does the intelligence of an AI-enabled city reside?

For the Technological City, intelligence resides in increasingly capable computational systems. For the Human-Centred City, it resides in the capacity to recognize and respond to heterogeneous human needs. For the Generative Planning City, it lies in the ability to participate in the production and reasoning of urban futures. For the Operational and Governable City, it lies in coordinated management and urban operating capacity. For the Complexity-Efficient City, intelligence reduces the cost of urban complexity. For the Adaptive and Iterative City, it resides in continuous feedback, learning, and revision. For the Hybrid Intelligence City, it emerges from relationships among human, artificial, institutional, collective, and embodied intelligences.

The central thesis of this article is that none of these schools is sufficient on its own. Urban intelligence is becoming relational and emergent. It does not reside exclusively in one model, a control room, a planner, an institution, a robot, or even a multi-agent system. It forms through continuous interactions among heterogeneous agents, people, infrastructures, institutions, and environments. We use the term CityAI to describe this emerging scale of intelligence.

CityAI is not a synonym for AI for cities. Nor does it imply a single super-intelligence controlling an urban system. CityAI refers to the emergent capacity of an urban system to perceive, interpret, reason, coordinate, act, evaluate, and evolve through the continuous interaction of heterogeneous human and artificial agents. This distinction matters. “AI for the city” starts with a technology and searches for urban applications. CityAI starts with the city as a complex, dynamic, multi-actor system and asks how urban-scale intelligence can form.

The question is therefore no longer only how AI can be used in cities. The more consequential step is whether intelligence itself is becoming a new form of urban infrastructure. If the twentieth-century city was transformed by transportation, electricity, and communication networks, the AI-era city may be transformed by infrastructures of perception, semantics, reasoning, coordination, and learning. Understanding this transition requires a historical perspective, which is developed in the next section.

1.2 Research approach and analytical procedure

This article adopts a historical-conceptual review and interpretive synthesis rather than a systematic bibliometric review. Its purpose is not to establish statistically exhaustive schools of urban AI, but to identify recurring intellectual positions through which urban intelligence has been conceptualized across urban planning, urban computing, smart-city research, Digital Twins, artificial intelligence, multi-agent systems, agentic AI, human-centred AI, and urban governance.

The interpretive procedure followed three steps. First, we examined major bodies of literature that have shaped the relationship between artificial intelligence and cities. Representative works were selected purposively on the basis of their conceptual influence, relevance to urban intelligence, and ability to articulate a distinctive interpretation of what makes a city intelligent.

Second, the literature was compared through three analytical questions: where does urban intelligence reside; what is that intelligence expected to achieve; and what limitation becomes visible when the perspective is considered in isolation? A vision was retained when a recurring body of work could be distinguished through a characteristic locus of intelligence, primary objective, and corresponding limitation. Representative scholars and works were associated with a vision according to their dominant analytical contribution, while recognizing that individual scholars and publications may contribute to several visions.

Third, boundaries between visions were defined analytically rather than disciplinarily. For example, the Operational and Governable City locates intelligence in real-time coordination, institutional authority, and urban operating capacity, whereas the Complexity-Efficient City focuses on reducing the information, coordination, management, response, and trial-and-error costs generated by urban interdependence. Similarly, the Adaptive and Iterative City emphasizes temporal feedback and continuous revision, whereas the Hybrid Intelligence City focuses on relationships among heterogeneous forms of human, artificial, institutional, collective, and embodied intelligence.

The seven visions should therefore be understood as an interpretive typology rather than as mutually exclusive or bibliometrically established schools. Rather than claiming that they have empirically converged, we argue more cautiously that they are becoming increasingly interdependent.

Accordingly, the originality of this article lies not only in introducing new terminology, but in developing an analytical framework for organizing, distinguishing, and evaluating emerging forms of urban intelligence. Specifically, the article contributes (1) a historical-conceptual account of the changing unit and capability of urban intelligence, (2) a seven-vision analytical typology, and (3) an operational framework that provides a basis for future empirical assessment of CityAI rather than treating it only as a definitional proposition.

2 History: From the Computed City to the Agentic City

The history of urban intelligence should not be written as a simple chronology of increasingly powerful computers. Such a narrative risks technological determinism and overlooks changes in how the city itself has been conceptualized [3]. A more useful history follows the changing relationship between computation and the urban object. Across roughly seven decades, the city has successively been treated as something to calculate, simulate, sense, learn, and increasingly to endow with distributed agency. These phases overlap, and older paradigms continue to coexist with newer ones. Nevertheless, they reveal a clear expansion in the assumed unit and capacity of urban intelligence (Fig.1).

2.1 The computed city: 1950s–1970s

The first major phase can be described as the computed city. After the Second World War, operations research, cybernetics, systems analysis, and early computer modelling created new confidence in the possibility of mathematically representing complex social and spatial systems [1,2,3]. Norbert Wiener’s cybernetics placed communication, control, and feedback at the centre of a general science of systems. Herbert Simon’s work on bounded rationality and complex systems challenged the assumption of perfectly informed decision-makers while simultaneously opening new ways to formalize decision processes. In planning and regional science, computational methods were increasingly used to represent land use, transport, population, and economic interactions.

Jay W. Forrester’s urban dynamics is emblematic of this period [3]. The city was represented through stocks, flows, feedback loops, and system-level relationships. The ambition was not to reproduce every individual urban actor. Instead, it was to identify the structural dynamics that generated long-term urban behaviour. Large-scale urban models similarly attempted to integrate land use and transportation through formalized relationships. Britton Harris and other pioneers of urban modelling helped establish the idea that planning can be supported by computational representations of urban systems models.

The intellectual image of the city in this era was the system. Its intelligence was external. Experts observed the city, formulated variables and equations, ran models, and interpreted results. Computation expanded the analytical reach of planners, but the model itself did not perceive the city continuously or act within it. Data were scarce, computational resources were limited, and calibration was difficult. Moreover, critiques of comprehensive modelling exposed the political and epistemological limitations of treating cities as closed, controllable systems.

Yet the computed city established several enduring ideas. First, urban processes can be represented computationally. Second, feedback mattered. Third, interventions can have indirect and delayed consequences. Fourth, the behaviour of the whole system can differ from the apparent logic of individual components. These ideas later became central to complexity science, Digital Twins, and adaptive urban systems.

2.2 The simulated city: 1970s–1990s

The second phase shifted attention from aggregate equations toward decentralized interaction. Complexity theory, cellular automata, artificial life, and agent-based modelling offered a different way to understand cities [16,17,26–29]. Rather than assuming that urban patterns must be specified at the macro level, researchers asked whether large-scale structures can emerge from local rules and interactions.

Thomas Schelling’s segregation model became a classic demonstration of emergence. Even relatively mild individual preferences can produce strongly segregated aggregate patterns. The significance of the model extended far beyond segregation. It showed that the relationship between individual intention and collective outcome can be nonlinear and counterintuitive. Urban patterns might be generated rather than centrally designed.

Cellular automata provided another influential framework. Space was divided into cells whose states changed according to local transition rules [30–32]. These models became widely used in urban growth and land-use change research. Michael Batty’s work on cities and complexity helped connect these computational approaches to a broader understanding of cities as evolving, self-organizing systems [33]. John Holland’s work on complex adaptive systems and Joshua Epstein and Robert Axtell’s artificial societies further advanced the idea that heterogeneous agents following rules can generate population-level phenomena.

The simulated city therefore introduced a crucial shift: from equations describing cities to agents generating cities. The unit of analysis moved downward, while the object of explanation remained collective. This was the intellectual foundation of contemporary multi-agent urban simulation.

Nevertheless, the simulated city introduced the idea that urban intelligence might be distributed. No single agent needed to understand the entire city for complex urban patterns to emerge. This insight remains essential for CityAI. The difference is that contemporary agents can increasingly perceive richer contexts, communicate in natural language, maintain memory, and invoke external tools. The question is no longer whether agents can generate urban patterns, but what happens when agents with increasingly rich cognitive capabilities interact within urban models.

2.3 The smart and sensed city: 1990s–2010s

The third phase was driven by digital networks, geographic information systems, mobile devices, ubiquitous computing, and the Internet of Things. The city became increasingly equipped with sensors, enabling continuous observation of urban processes [4,34–36]. The smart-city paradigm emerged as governments and technology firms connected digital infrastructure to urban management and service delivery.

William J. Mitchell’s work on digital urbanism examined the changing relationship between networked information and physical space. Carlo Ratti and the Senseable City approach demonstrated how real-time data can reveal previously invisible urban dynamics. Anthony Townsend critically examined the promises and politics of smart cities. In urban computing, researchers such as Yu Zheng developed methods for extracting knowledge from heterogeneous urban data, including mobility trajectories, points of interest, environmental observations, and social data.

The decisive change was temporal. Earlier urban models often relied on censuses, surveys, and periodic data collection [5–7]. The sensed city can be observed continuously or at much higher frequencies. Traffic speed, air quality, energy consumption, mobile activity, and human movement can be transformed into streams [8,37]. The city became computationally visible in near real time.

This transformation enabled a new class of urban applications. Traffic control systems can respond to changing traffic flows. Environmental monitoring can identify local anomalies. Location-based services can adapt to individual movement. Urban dashboards can integrate multiple data sources. The smart city became a platform for measurement and operational awareness.

But most importantly, sensing does not equal intelligence. A city can collect enormous quantities of data while remaining institutionally fragmented and operationally slow. Dashboards can display conditions without generating coordinated responses. The smart-city era made the city observable, but it did not fully solve the challenges of semantic interpretation and agency.

2.4 The learning city: 2010s–2020

The expansion of Machine Learning and Deep Learning produced the fourth phase: the learning city [11]. Urban analytics increasingly shifted from manually specified relationships toward models capable of learning complex patterns from data.

Machine Learning was applied to traffic forecasting, building energy demand, air quality, land-use classification, disaster risk, crime prediction, mobility, and urban morphology. Convolutional neural networks transformed the analysis of street-view and remote-sensing imagery [4,11]. Recurrent and graph neural networks expanded spatiotemporal modelling. Gradient boosting methods became powerful tools for urban prediction and feature interpretation. The rise of large urban datasets made it possible to model relationships at scales that were previously difficult to observe.

The conceptual promise was significant. The city no longer needed to be fully specified through equations before computation can begin. Models can learn latent relationships from examples. In some fields, prediction became dramatically more accurate.

However, the learning city was predominantly predictive. A model estimated a future traffic state, classified an image, predicted energy demand, or identified risk. Even when outputs supported decisions, the model usually did not formulate objectives, negotiate trade-offs, or act autonomously. Moreover, high predictive performance can conceal weaknesses in causal interpretation, transferability, fairness, and explainability.

The learning city also exposed a scale problem. Urban AI research became highly specialized. One model predicted traffic, another one estimated heat exposure, another detected land-use change, and another one optimized energy. Each system can be intelligent within a narrow task while the city as a whole remained fragmented. Technical intelligence increased faster than the systemic intelligence.

2.5 The generative and agentic city: 2020–2025

The emergence of foundation models and Large Language Models initiated another transition. The importance of these models lies not only in text generation [12−14]. Their broader significance is the combination of general-purpose semantics, contextual interpretation, reasoning-like capabilities, multimodal processing, tool use and external data integration.

Urban foundation models have been proposed as large-scale models pre-trained on heterogeneous, multi-source, and multi-granularity urban data. Their ambition is to develop transferable capabilities across urban domains rather than training a separate model for every task. Zhang et al. [62] explicitly connect this agenda to Urban General Intelligence: not artificial general intelligence in a philosophical sense, but a research direction toward more versatile and transferable urban intelligence.

At the same time, generative AI entered urban planning and design. Models began to generate spatial configurations, design alternatives, images, text, codes, and scenarios [14,18,38]. The relationship between AI and planning shifted from analysis to production [39]. AI can participate in making models of possible urban futures.

The agent paradigm extended this shift. A large language model can be embedded in an architecture with sensing, memory, goals, planning routines, external tools, and action interfaces. Multiple agents can communicate, divide tasks, critique one another, or represent different stakeholders. Recent concepts such as intelligent urban agents and Agentic Urban AI explicitly describe a movement beyond human-directed automation toward systems capable of dynamic goal adjustment, strategic adaptation, and multi-agent coordination [68,69].

This development reconnects contemporary AI with the older tradition of agent-based modelling, but the nature of the agent has fundamentally changed—from acting in closed worlds to acting in open worlds. A classical urban agent might follow a rule such as “move if neighbourhood satisfaction falls below a threshold”. An LLM-powered urban agent may receive multimodal observations, interpret a policy document, recall prior events, communicate with other agents, invoke a mobility model, and generate a contextual response [20–22]. The gain is expressive and semantic capacity [40,23]. The cost is greater uncertainty, opacity, computational expense, and evaluation difficulty.

The generative and agentic city therefore marks a movement from models that describe urban systems toward computational entities that participate in them. AI can increasingly function as a planner, negotiator, simulated citizen, service agent, or embodied robot, but these roles remain bounded by institutional, legal, and human constraints. Rather than disappearing, the boundary between urban models and urban actors is progressively being redefined.

2.6 Toward the CityAI era: 2025 and beyond

The current transition can be understood as another change in the unit of intelligence. The early computed city centred on the model [41]. The learning city centred on task-specific predictive systems. Foundation models expanded the generality of the model. Agents connected models to memory, tools, objectives, and actions. Multi-agent systems connected agents to one another in cooperation or competition settings. The next conceptual step is to ask whether an ecosystem of heterogeneous intelligences can generate capabilities at the scale of the urban system.

The sequence is no longer simply one of model, larger model, and even larger model. It is object intelligence, individual intelligence, collective intelligence, and potentially urban intelligence. In object intelligence, a device or model performs a bounded task. In individual intelligence, an agent integrates perception, reasoning, and action around goals. In collective intelligence, multiple agents coordinate or compete. In urban intelligence, heterogeneous human and artificial actors interact with infrastructures, institutions, and environments through continuous feedback.

The historical trajectory can be summarized as follows: the city was first calculated, then simulated, sensed, learned, generated, and increasingly inhabited by artificial agents. Each transition expanded the relationship between computation and urban life. The emerging question is whether these capacities can be integrated without collapsing urban plurality into a single optimization objective. This question is at the centre of the seven contemporary visions of the AI City.

3 Seven Global Visions for the AI City

The contemporary field of urban AI is diverse enough that a single technological taxonomy is insufficient. Classifying research by Machine Learning, computer vision, Digital Twins, Large Language Models, or robotics describes the tools but not the underlying intellectual ambitions. Likewise, classifying work by transportation, energy, planning, governance, or public safety describes application domains but does not reveal how different communities understand urban intelligence.

This section proposes seven visions. They are analytical constructs derived from recurring orientations in research and practice. The purpose is not to force every scholar into one category. Rather, the seven visions reveal different answers to three questions: what is the unit of urban intelligence? What is intelligence expected to achieve? What limitations become visible when that vision is applied alone? (Table 1)

3.1 Vision one: the technological city—the AI city as a stack of emerging technologies

The Technological City is the most visible and rapidly changing vision. Its central question is: what technical capabilities are required to make a city intelligent? [11–13] Progress is understood through the accumulation and integration of increasingly advanced technologies.

At least twelve technological families now shape the AI City discourse: Internet of Things and ubiquitous sensing; remote sensing and geospatial intelligence; Machine Learning; Deep Learning; computer vision; natural language processing; knowledge graphs and semantic technologies; generative AI; foundation models; Digital Twins; multi-agent and agentic AI; and embodied intelligence and robotics [14,15]. Edge AI, federated learning, graph learning, reinforcement learning, and privacy-enhancing computation further extend this landscape.

These technologies can be interpreted not simply as a list but as an expanding capability stack. Sensors allow the city to sense (e.g., cameras, LiDAR, IoT devices...). Computer vision allows it to see. Machine Learning allows it to learn patterns. Knowledge graphs allow entities and relationships to be structured. Large Language Models provide semantic interpretation and natural-language interaction. Generative AI enables the production of alternatives. Digital Twins provide representational and experimental environments. Agents connect models to goals, memory, tools, and action. Multi-agent systems enable coordination and competition. Robotics and autonomous systems embody computational decisions in physical space.

This capability-based interpretation is important because it shows why the recent technological convergence is different from the earlier digitization. A traditional smart-city platform might integrate data from multiple departments [20,22,40]. An AI City stack potentially connects perception, semantic interpretation, prediction, reasoning, simulation, coordination, and physical response. The system is no longer only a data pipeline [42–44]. It begins to resemble a cognitive-action loop.

The Technological City is associated with several research communities. Urban computing has developed methods for acquiring, integrating, and mining heterogeneous urban data. GIScience has advanced geospatial representation and spatial AI. Computer vision has transformed the interpretation of street-level and remote-sensing imagery. Urban foundation model research seeks generalizable representations across multiple data modalities and urban tasks [62]. Agentic AI research connects large models to planning, tools, and actions. Digital twin research integrates real-time data with dynamic models and scenario testing.

Several intellectual lineages are visible here. Batty’s computational urban science established a powerful account of cities as complex and computable systems; the Senseable City agenda foregrounded real-time sensing and feedback between digital information and urban life; and Zheng and colleagues systematized urban computing as knowledge discovery from heterogeneous urban data [4,27]. More recently, the Urban Foundation Model literature has explicitly framed cross-task and cross-modality transfer as a route toward Urban General Intelligence. Reasoning-oriented planning research has simultaneously argued that statistical learning alone is insufficient for value-based, rule-grounded, and explainable planning decisions [66].

The strength of the Technological City is capability expansion. It continually asks what urban systems can do that was previously impossible. Without this vision, CityAI would drop its computational substrate.

The Technological City therefore answers the question “with what can urban intelligence be built?” It does not fully answer “for whom, toward what values, or through what collective process?” These questions lead to the second vision.

3.2 Vision two: the Human-Centred city—the AI city as a response to heterogeneous human needs

The Human-Centred City begins with a different question: intelligence for whom? It challenges the assumption that urban intelligence can be measured by computational capacity, automation rate, or system efficiency alone [36,45,46]. A city is intelligent to the extent that it can recognize, interpret, and respond to diverse human needs while protecting rights and enabling human flourishing.

This vision has roots in human-centred design, participatory planning, inclusive smart cities, responsible AI, and people-centred digital transformation. International urban AI governance frameworks increasingly emphasize that AI adoption should serve residents rather than technological experimentation for its own sake. Responsible AI in cities raises issues of fairness, transparency, privacy, accountability, accessibility, and public participation.

The key intellectual contribution of this vision is heterogeneity. There is no average urban human [7,8,37]. The same environmental condition can produce radically different experiences for a child, an older adult, a wheelchair user, an outdoor worker, a tourist, or a person with chronic vulnerability. The same transport disruption affects a high-income remote worker differently from a low-income shift worker. The same digital public service may empower one resident while excluding another because of language, disability, connectivity, or digital literacy.

Human needs can be described through multiple frameworks, but for the AI City, they may be grouped into at least several recurring dimensions: safety, health, thermal and environmental comfort, mobility and accessibility, social connection, economic opportunity, identity and belonging, dignity, fairness, autonomy, and long-term well-being. These dimensions are not independent. A mobility intervention can improve accessibility but reduce neighbourhood social cohesion. A security system can increase perceived safety for some groups while increasing surveillance anxiety for others. An efficiency-oriented service may disadvantage people whose behaviour does not fit the dominant data pattern.

AI creates new opportunities for recognizing heterogeneity. Multi-source sensing can reveal local environmental exposure. Computer vision and mobility data can identify spatial barriers. AI translation systems can support multilingual interaction. Personalized agents can adapt information to user context. Large-scale simulations can represent populations with differentiated ages, occupations, incomes, preferences, vulnerabilities, and social networks. Multi-agent systems can model conflict among stakeholder objectives rather than collapsing them into a single average utility.

However, computationally representing heterogeneity introduces new risks. A persona is not a person [45,47,10]. A synthetic population can reproduce biases in source data. An LLM agent representing an older adult or minority community may generate plausible language without faithfully representing lived experience [48]. Human-centred CityAI therefore requires a distinction between computational representation and political representation. Simulating a stakeholder cannot replace giving that stakeholder institutional voice.

The Human-Centred City also changes the evaluation of urban intelligence. Instead of asking only whether a model is accurate, it asks whether a system improves outcomes across different groups. Instead of average travel time, it may examine accessibility for vulnerable populations. Instead of aggregate thermal comfort, it may examine exposure inequalities. Instead of overall service efficiency, it may evaluate who is systematically excluded.

The representative intellectual lineage includes Jane Jacobs’s attention to lived urban complexity [63], Jan Gehl’s human-scale approach to urban life and public space [64], Amartya Sen’s capability approach as applied to development and well-being, human-centred smart-city scholarship, and contemporary responsible AI and participatory technology research. In the AI context, the emphasis on collective intelligence and democratic input developed by scholars and organizations concerned with AI governance also becomes relevant. The core proposition is that urban intelligence must be evaluated through the plurality of urban lives.

This challenge directly connects the Human-Centred City to multi-agent intelligence. Agents may help represent, explore, and deliberate over diverse perspectives. But the purpose is not to create an artificial substitute for society. It is to increase the capacity of urban systems to perceive difference, expose trade-offs, and support more inclusive human decisions.

3.3 Vision three: the Generative Planning City—AI as planner and designer

The third vision asks whether AI can participate in making the city. Planning has historically used computation primarily for analysis: population forecasting, transport modelling, land use, environmental assessment, and scenario evaluation [49]. The rise of Generative and Agentic AI expands AI’s role from analysing plans to producing and reasoning about planning alternatives.

This transition can be described in three stages. The first is AI for planning analysis. Machine Learning identifies patterns, predicts conditions, or estimates impacts. The planner remains the primary generator of alternatives. The second is AI for planning generation. Generative models, optimization algorithms, reinforcement learning, and procedural systems produce land-use configurations, building layouts, street networks, images, or design options. The third is AI for planning reasoning. Agents interpret objectives and regulations, use professional tools, compare alternatives, verify constraints, deliberate over stakeholder values, and explain recommendations.

The third stage is particularly important. Urban planning is not a pure optimization problem because planning objectives are normative, contested, and institutionally constrained [14,50]. A statistically likely land-use configuration is not necessarily a desirable one. A visually convincing design may violate regulations. A mathematically efficient solution may be politically unacceptable or socially unjust. Planning requires explicit engagement with rules, values, uncertainty, and trade-offs.

Recent agentic planning research reflects this problem. Qian et al. [67] used consensus-based multi-agent reinforcement learning to represent stakeholder preferences in land-use readjustment. Yang, Li, and Biljecki [66] propose a reasoning-oriented planning framework combining perception, domain foundations, and explicit functions of analysis, generation, verification, evaluation, collaboration, and decision. Together, these studies indicate a shift from AI that predicts urban conditions toward AI that participates in structured planning reasoning.

The intellectual shift is from computer-aided planning to computational participation in planning cognition. AI becomes capable of proposing, critiquing, checking, and revising [49,51,39]. This does not imply the disappearance of the planner. On the contrary, it will increase the importance of professional judgement because the solution space expands dramatically. The planner’s role shifts from manually producing every alternative toward framing objectives, curating evidence, defining constraints, evaluating values, and governing human–AI collaboration.

Representative figures in computational planning include Michael Batty, whose work on complexity and computation reshaped urban modelling; researchers in generative urban design and spatial AI; and emerging scholars working on agentic planning. Filip Biljecki and collaborators have recently argued that planning AI must move beyond statistical learning toward value-based, rule-grounded, and explainable reasoning. Research on AI agents as urban planners similarly explores collective decision-making and stakeholder dynamics.

The Generative Planning City also introduces a new temporal opportunity. Conventional plans are expensive to produce and therefore often updated slowly. If AI can continuously integrate new data, regenerate scenarios, and evaluate changing conditions, planning can become more iterative. The boundary between planning and operation begins to dissolve. A plan could become a living hypothesis rather than a fixed document.

The Generative Planning City locates intelligence in the capacity to produce and reason about possible urban futures. Its limitation is that planning outputs, however intelligent, remain incomplete if the city cannot continuously operate, evaluate, and adapt. This leads to the operational vision.

3.4 Vision four: the Operational and Governable City—AI as urban operating capacity

The Operational and Governable City understands intelligence as the capacity to coordinate urban systems in real time [20,21,40,43,52]. Its central question is: can a city operate itself better?

This vision is visible in traffic control, energy management, water systems, waste collection, emergency response, infrastructure maintenance, public safety, and digital public services. It is associated with city brains, urban operating systems, command centres, integrated operations platforms, and AI-supported governance.

The metaphor of the city as an operating system is powerful because it focuses attention on coordination. Urban problems frequently cross administrative and infrastructural boundaries [53]. A heatwave is simultaneously an environmental, health, energy, mobility, and social vulnerability problem. A flood affects roads, emergency services, electricity, communication, schools, hospitals, and public behaviour. Departmental intelligence is insufficient when the event itself is cross-systemic.

AI can increase operational capacity in several ways. Predictive models anticipate demand or failure. Optimization allocates resources. Computer vision detects anomalies. Natural-language systems summarize information and support administrative workflows. Digital Twins provide common situational representations. Agents can monitor subsystems, invoke tools, and coordinate tasks. Multi-agent systems can distribute responsibilities across specialized agents while maintaining communication and coordination.

The AI City framework’s urban brain, cerebellum, and central nervous system can be interpreted within this operational lineage. The brain concerns higher-level intelligence and decision support; the cerebellum concerns rapid coordination and execution; the nervous system concerns distributed perception and communication [19,46,54,48]. This biological metaphor is not a literal claim that a city is a living body. Its value lies in distinguishing functions that conventional smart-city architectures often collapse into a single “platform”.

City Brain initiatives provide another operational model. Alibaba’s City Brain, initially developed in Hangzhou, became widely associated with AI-supported traffic management and the integration of urban data for operational decision-making. Singapore’s Smart Nation and digital twin initiatives, Dubai’s digital governance ambitions, and integrated command platforms in many cities similarly reflect the desire for real-time urban management.

International organizations increasingly frame AI applications around urban planning, management, and governance. Organisation for Economic Co-operation and Development’s (OECD) 2025 issues note identifies mobility, housing, infrastructure, public services, and urban management as major areas of AI adoption, while UN-Habitat emphasizes the governance structures and local practices required for responsible urban AI [54,65]. Mobility and utilities are especially prominent because they contain continuous flows, measurable performance indicators, and operational control points.

However, the Operational City exposes a crucial distinction: optimization capabilities are not intelligence. A city that operates efficiently is not necessarily a city that understands. Operational systems may optimize the variables they can measure while neglecting values they cannot easily quantify. Centralized platforms may increase coordination but also create surveillance, cybersecurity, and power-concentration risks. Automated decisions may become difficult to contest. Fast response can amplify a bad objective as effectively as a good one.

Governability is therefore as important as operation. An AI-enabled city must specify authority, permissions, accountability, auditability, and escalation. Which agent may only observe? Which may recommend? Which may act? Under what conditions must a human approve? Which actions are reversible? How are conflicting departmental objectives resolved? How are residents informed and given channels for contestation?

The future Operational City will likely be multi-agent because urban functions are inherently distributed. Yet the design of agent roles should reflect institutional structures and legal responsibility rather than only software convenience. A transport agent, emergency agent, energy agent, and public communication agent may coordinate during a crisis, but their authority cannot be invented by the model. CityAI requires an institutional semantics of action.

3.5 Vision five: the Complexity-Efficient City—AI as a means of reducing the cost of urban complexity

The fifth vision is often present but rarely named as an independent intellectual perspective [9,47,10]. It asks: can AI reduce the cost of coordinating urban complexity?

Cities generate enormous coordination costs. As urban scale and interdependence increase, information must travel across more organizations, infrastructures, and social groups. Decisions require more data. Maintenance networks become larger. Service demand becomes more heterogeneous. The consequences of error become more expensive. We propose the term “cost of urban complexity” to describe the combined information, coordination, management, response, and trial-and-error costs generated by increasing urban interdependence.

The Complexity-Efficient City does not reduce intelligence to financial savings. Rather, it interprets intelligence as the capacity to achieve better coordination with lower marginal cost [4,6,9]. This perspective is particularly relevant to public administrations facing limited budgets, infrastructure owners managing large asset portfolios, and rapidly growing cities that cannot expand administrative capacity at the same rate as urban complexity.

AI can reduce information costs by automatically processing large data streams. It can reduce inspection costs through computer vision and remote sensing. Predictive maintenance can shift infrastructure management from reactive repair toward risk-based intervention. Optimization can improve resource allocation. Generative AI can reduce the time required to produce routine documents or alternatives. Digital Twins can allow scenarios to be tested before physical implementation. Agents can automate parts of cross-system coordination.

The economic value of simulation is especially important. Urban interventions are expensive and often irreversible. A bridge, metro line, land-use plan, or flood-control system cannot be tested repeatedly in the physical city. Simulation creates a lower-cost experimental space. As urban models become more dynamic and agentic, they may enable not only engineering scenario testing but also the exploration of behavioural and institutional responses.

This leads to a broader interpretation of AI as a technology for reducing the cost of urban trial and error. The traditional city learns slowly because feedback is delayed and interventions are costly. A CityAI system may accelerate learning by generating hypotheses, simulating alternatives, identifying uncertainty, and monitoring real-world outcomes after implementation.

However, the Complexity-Efficient City has a serious limitation: value reduction. What appears as a cost in one accounting system may be a source of social value in another. Redundancy can look inefficient but provide resilience. A small public facility may have low utilization but be essential to a vulnerable community. Human deliberation may be slower than automation but generate legitimacy. Informal urban practices may resist standardization while supporting social networks.

The central danger is to confuse lower operational cost with higher urban intelligence. CityAI must distinguish efficiency from value. The purpose of reducing the cost of complexity is not to eliminate complexity. Cities are complex partly because human needs, identities, and institutions are plural. The goal is to increase the capacity to coordinate complexity without erasing it.

3.6 Vision six: the Adaptive and Iterative City — the AI city as a continuously evolving system

The Adaptive and Iterative City asks perhaps the most fundamental question for urban intelligence: can a city learn from its own operation [1,3,27,28]?

Conventional planning is often represented as a sequence: plan, build, use. Although real planning systems include monitoring and revision, the practical cycle can be slow. Plans are produced through lengthy procedures, physical interventions are implemented, and formal updates may occur much later. This temporal structure is increasingly misaligned with rapid environmental, technological, demographic, and behavioural change.

The adaptive vision replaces the linear sequence with a loop: sense, interpret, decide, act, evaluate, learn, and adapt [16,17,30,31,56]. Intelligence resides not in finding a permanently optimal answer but in maintaining the capacity to revise an answer as conditions change.

This perspective has deep roots in cybernetics, complex adaptive systems, resilience thinking, and adaptive management. Contemporary AI gives the loop new technical possibilities. Ubiquitous sensing increases feedback frequency. Machine Learning detects patterns. Digital Twins create experimental representations. Generative models produce alternatives. Agents coordinate actions. Evaluation systems compare expected and observed outcomes. Learning mechanisms update models or policies.

The AI City proposition explicitly emphasized an iteration mode [19]. This differentiates AI City thinking from static digital infrastructure [49,51,32]. More recent Agentic Urban AI research similarly highlights dynamic goal adaptation and strategic adjustment as characteristics that distinguish agentic systems from earlier reactive smart city technologies [69].

For AI-driven adaptive urbanism, adaptation must be understood at multiple scales. At the micro scale, a public space may change lighting, cooling, information, or robotic services in response to environmental conditions and human behaviour. At the district scale, mobility services may adjust routes or resource allocation. At the city scale, planning strategies may be revised as climate risks, population distributions, or economic structures change. At the institutional scale, governance protocols themselves may evolve in response to observed performance.

The adaptive city also requires a theory of memory. A system that reacts to current conditions without retaining experience is responsive but not necessarily learning. CityAI must distinguish immediate reaction, short-term adaptation, and long-term institutional learning. Urban memory may be distributed across databases, models, planning documents, laws, agent memories, and human organizations. Connecting these forms of memory is a major challenge.

Another challenge is evaluation. To learn, the system must compare action and outcome. But urban outcomes are multicausal and delayed. If a heat intervention is followed by lower health risk, was the intervention responsible, or did weather change? If a public-space adaptation attracts more people and subsequently increases heat exposure, was the intervention successful or counterproductive? Adaptive CityAI therefore requires causal reasoning, uncertainty estimation, and counterfactual simulation rather than simple feedback optimization.

The adaptive vision also changes the meaning of planning. A plan becomes less like a final image and more like a policy for continuous adjustment. This does not mean that long-term visions disappear. On the contrary, adaptation without direction can become opportunistic. The city needs relatively stable values and goals alongside flexible strategies. The distinction between what should remain stable and what should adapt is itself a planning question.

The representative intellectual lineage includes Norbert Wiener’s feedback systems, John Holland’s complex adaptive systems, resilience and adaptive management scholarship, and contemporary research on adaptive cities and agentic AI. The AI-driven Adaptive Urbanism perspective extends this lineage by asking how AI can discover urban mechanisms, support iterative intervention, and connect sensing, simulation, action, and learning.

Its strength is temporal intelligence. It recognizes that urban conditions and human responses change. Its limitation is coordination. Multiple adaptive agents can produce instability if each one optimizes locally. A traffic system adapts, residents adapt to the traffic system, land values adapt to accessibility, developers adapt to land values, and planning regulations adapt to development pressure. Adaptation is recursive.

3.7 Vision seven: the Hybrid Intelligence City—the city as a habitat of human and artificial intelligence

The Hybrid Intelligence City begins with the recognition that the future city will not be controlled by one AI [18,38,50]. It will be inhabited by many intelligences.

These intelligences include human intelligence, Artificial Intelligence, collective intelligence, institutional intelligence, and embodied intelligence. Residents perceive and interpret urban conditions through lived experience. Communities generate collective knowledge. Governments hold institutional memory and legitimate authority. AI models identify patterns and generate recommendations. Digital agents reason, communicate, and use tools. Robots and autonomous vehicles act physically. Buildings and infrastructures may contain local control systems. The city becomes a habitat in which different forms of intelligence coexist and interact.

Hybrid intelligence research argues that human and machine capabilities can be combined in socio-technical ensembles through complementarity and mutual learning [18]. In the urban context, this idea must be expanded [45,46]. The relevant ensemble is not a small human–AI team. It may contain millions of people and large populations of artificial agents distributed across a cyber-physical-social environment.

The central question is therefore not “human or AI?” It is “what relationships among heterogeneous intelligences produce legitimate, robust, and adaptive urban capacities?”

This perspective rejects two simplistic futures. The first is full automation: the idea that a sufficiently powerful central AI can optimize the city. The second is AI as a passive tool: the idea that artificial systems will remain isolated instruments used only when humans explicitly request them. Urban reality is likely to be more entangled. AI recommendations will change human behaviour. Human behaviour will generate new data. Institutions will regulate AI. AI will reshape institutional workflows. Robots will alter public-space expectations. New urban forms may be designed around autonomous systems. The changed city will then alter both human and artificial behaviour.

This is co-adaptation. Humans change AI; AI changes human behaviour; both change the city; and the changed city changes both again.

Multi-agent systems provide a partial technical model for this condition. Different agents can possess specialized roles, goals, memories, and tools. They can communicate and coordinate. However, urban hybrid intelligence exceeds conventional multi-agent AI in three ways.

First, urban agents are radically heterogeneous. A resident, a planning institution, a drainage system, a delivery robot, and a language-model agent do not share the same ontology, timescale, or action space.

Second, power is asymmetric. Some agents can recommend, some can purchase, some can regulate, and some can physically intervene [14,18,38]. Intelligence cannot be separated from authority.

Third, the urban environment is not a neutral simulation space. It contains physical limits, historical inequalities, legal rights, ecological thresholds, and cultural meanings.

We therefore propose the concept of the Ultra-Agent Habitat. An Ultra-Agent Habitat is an urban environment in which heterogeneous human and artificial agents continuously perceive, communicate, negotiate, coordinate, act, and co-evolve through shared physical, digital, and semantic infrastructures. The word “habitat” is important. It shifts attention from the architecture of a single intelligent system to the conditions that allow many intelligences to coexist.

The Ultra-Agent Habitat is related to, but analytically distinct from, established concepts such as large socio-technical systems and cyber-physical-social systems. These frameworks primarily describe the coupling of social, technical, and physical components. The Ultra-Agent Habitat places additional emphasis on heterogeneous agency: different entities may possess different degrees of perception, memory, reasoning, autonomy, authority, and capacity to act. These capacities are neither uniform nor symmetrical.

The concept is also explicitly open-world and co-evolutionary. Its participants do not interact only within a predefined computational architecture. Human actors, artificial agents, institutions, infrastructures, and environments continuously modify one another, while operating across different timescales, action spaces, legal constraints, and institutional responsibilities. The central concern is therefore not simply system integration, but the conditions under which heterogeneous intelligences can coexist, coordinate, contest, and evolve within a shared urban habitat.

The Ultra-Agent Habitat requires more than connectivity. It requires shared semantics. A sensor may report 34 °C. A human may say “I feel dizzy”. A camera may detect reduced walking speed. A health protocol may define heat-risk thresholds. A robot may possess water-delivery and navigation capabilities. These are different data and action languages. Coordination requires a semantic layer capable of translating heterogeneous signals into a shared urban state without erasing their source and uncertainty.

This challenge points to the need for what we refer to as an Urban Interlingua: a shared semantic layer that enables humans, artificial agents, environments, and institutions to achieve sufficient mutual interpretability for coordinated action. Here, the term refers to a supporting mechanism for semantic interoperability within the Ultra-Agent Habitat, rather than an independent urban intelligence framework.

The Hybrid Intelligence City is the closest of the seven visions to CityAI because it relocates intelligence from components to relationships. Yet it also creates the most difficult governance questions. How are values encoded? How are conflicts resolved? Can artificial agents negotiate on behalf of people? How are permissions authenticated? What happens when agents use different foundation models and produce incompatible interpretations? How can the system remain open to innovation without becoming insecure or ungovernable?

The emergence of an Ultra-Agent Habitat intensifies questions of power. Critical smart-city scholarship has demonstrated that digital infrastructures redistribute not only information, but also visibility, decision-making capacity, and institutional authority. These concerns become more consequential when artificial agents can initiate or mediate actions rather than merely sense, predict, or recommend.

Ownership is therefore a constitutive question for CityAI. Data, models, semantic representations, agent identities, and interaction logs may be distributed across governments, infrastructure operators, technology providers, communities, and individuals. Operating the technical infrastructure should not automatically confer authority over the shared semantic layer. Governance arrangements must specify who may create, modify, validate, and retire semantic entities and rules.

Auditability and authority must likewise be explicit. Agent actions should be traceable to identifiable evidence, models, permissions, and responsible institutions. Observing an urban condition, recommending an intervention, making an administrative decision, and physically acting through infrastructure should require different levels of authorization.

Contestation is equally fundamental. Urban interests are inherently plural, and conflicts cannot be resolved solely through optimization. When residents, infrastructure operators, governments, and private providers disagree, CityAI must preserve mechanisms for explanation, appeal, override, negotiation, and institutional review. Its purpose is not to eliminate political disagreement, but to make assumptions, authority, interactions, and consequences more visible and contestable.

3.8 From seven visions to CityAI

The seven visions constitute a fragmented but increasingly interdependent field. The Technological City locates intelligence in the model and technical stack [41]. Its goal is capability, and its limitation is fragmentation. The Human-Centred City locates intelligence in responsiveness to people. Its goal is well-being and inclusion, and its limitation is scaling representation without reduction. The Generative Planning City locates intelligence in planning cognition. Its goal is the production and reasoning of urban futures, and its limitation is the risk of static or unaccountable outputs. The Operational and Governable City locates intelligence in coordinated systems. Its goal is effective operation, and its limitation is optimization bias and centralization. The Complexity-Efficient City locates intelligence in coordination efficiency. Its goal is to reduce the cost of complexity, and its limitation is value reduction. The Adaptive and Iterative City locates intelligence in feedback loops. Its goal is continuous evolution, and its limitation is multi-system instability and coordination. The Hybrid Intelligence City locates intelligence in human–AI relationships. Its goal is collective and complementary intelligence, and its limitation is governance across heterogeneous agents.

The key theoretical finding is that all seven visions identify a real component of urban intelligence, but each becomes insufficient when isolated. Advanced technologies require human values. Human needs require scalable mechanisms of representation and response. Generative planning requires operational feedback. Operations require economic feasibility. Efficiency requires protection against value reduction. Adaptation requires coordination. Hybrid intelligence requires shared semantics and governance.

Earlier approaches tended to locate intelligence somewhere in the city. Intelligence was in algorithms, in a command platform, in the planner’s decision support system, in a responsive service, or in a feedback loop [20,57,23]. CityAI begins from a different proposition: urban intelligence does not reside in any single component [52,58,59]. It emerges from relationships among agents, people, infrastructures, institutions, and environments.

We define CityAI as the emergent capacity of an urban system to perceive, interpret, reason, coordinate, act, evaluate, and evolve through continuous interaction among heterogeneous human and artificial agents, infrastructures, institutions, and environments.

Seven verbs in this definition are deliberate: “perceive” means acquiring multimodal urban states. “Interpret” means transforming heterogeneous signals into contextual and semantic meaning. “Reason” means evaluating relationships, constraints, causes, values, and possible actions. “Coordinate” means aligning or negotiating among multiple agents and systems. “Act” means intervening through digital, institutional, or physical channels. “Evaluate” means comparing outcomes with objectives while recognizing uncertainty and distributional effects. “Evolve” means updating models, strategies, relationships, and, when legitimate, goals.

First, CityAI is not simply AI for cities. AI for cities is application-oriented. It begins with AI methods and applies them to transport, energy, planning, or other domains. CityAI is system-oriented. It asks how urban-scale intelligence is formed across domains.

Second, CityAI is not identical to a city brain. A city brain is usually an integrative or coordinating architecture. CityAI may include brain-like functions, but intelligence is distributed across human and artificial actors. No single centre possesses the complete urban state.

Third, CityAI is not equivalent to a multi-agent system. Multi-agent intelligence is a necessary precursor because coordination among autonomous entities is fundamental. But the city contains agents that are not software, goals that are contested, environments that are physical, and institutions that possess legitimate authority. CityAI is the urban-scale emergent capacity of this wider Ultra-Agent Habitat.

The conceptual progression can be understood as an expansion in the unit of analysis—from single models and agents, to multi-agent systems and agent ecosystems, and ultimately to urban-scale relational conditions such as the Ultra-Agent Habitat from which CityAI may emerge. Each step expands the unit of intelligence and the complexity of relationships that must be coordinated.

This perspective also changes the central research question. The challenge is no longer to build the one model that knows the city. Such a model is unlikely to be technically feasible or politically desirable. The challenge is to create conditions under which multiple intelligences can produce coherent urban capacities while preserving plurality, accountability, and adaptability.

4 Future: From Multi-Agent Intelligence to CityAI

The transition from multi-agent intelligence to CityAI will not be achieved simply by increasing the number of agents or improving the capability of individual models. It depends on whether heterogeneous human and artificial actors can generate coherent, adaptive, and governable capacities at the scale of the urban system. Multi-agent systems provide an important technical foundation because cities are inherently distributed across residents, institutions, infrastructures, environments, and operational subsystems. However, coordination among agents is not equivalent to urban intelligence. CityAI requires connections between micro-level behaviour, collective organization, urban reasoning, physical intervention, institutional authority, and long-term learning.

The following research agenda identifies five interrelated directions through which multi-agent intelligence may develop toward CityAI: scalability, collective urban intelligence, urban reasoning, living urban models, and human–AI co-evolution. These directions are not intended as independent technological trends, but as complementary capabilities that together provide the basis for the transition from the AI City to CityAI.

4.1 From small multi-agent systems to civilization-scale agent ecosystems

The first frontier is scale. Most agent-based urban models historically used simple agents because large populations were computationally expensive [16,17,30,31]. Contemporary Large Language Model agents reverse the problem: individual agents can display richer language and reasoning behaviours, but they are costly, slow, and difficult to evaluate at population scale.

CityAI requires research across a scale gradient: a few agents, thousands of agents, millions of agents, and potentially populations approaching the scale of large metropolitan regions. The scientific question is not merely how to run more agents. It is how intelligence changes with scale.

At small scales, researchers can examine detailed dialogue, memory, reasoning, and negotiation. At intermediate scales, social networks, group formation, information diffusion, and institutional interactions become important. At metropolitan scales, aggregate patterns, infrastructure loads, spatial inequality, and emergent behaviours become visible. A civilization-scale agent ecosystem may exhibit phenomena that cannot be inferred from small simulation.

This creates a methodological challenge. It is neither necessary nor computationally rational for every urban agent to use a large language model continuously [14]. Future CityAI simulations will likely use heterogeneous agent architectures. A small number of active agents may possess rich reasoning and dialogue. Larger populations may use compressed policies, behavioural models, probabilistic transitions, or learned surrogates. Agents may dynamically switch levels of cognition when events become relevant to them.

Such architectures can be described as variable-resolution intelligence. The simulation allocates cognitive resources where urban conditions require deeper reasoning. During normal conditions, most agents follow efficient behavioural models. During a flood, a household receiving conflicting warnings may activate a richer reasoning process. A community leader may communicate with an emergency agent. An infrastructure agent may invoke a hydraulic model. The system dynamically switches between different levels of reasoning at the appropriate knowledge granularity.

Evaluation must also change with scale. At the individual level, researchers may evaluate behavioural plausibility, consistency, memory, and reasoning. At the group level, they may examine consensus, polarization, information diffusion, and social network effects. At the urban level, they may evaluate accessibility, rescue efficiency, infrastructure peaks, emotional stability, distributional outcomes, and economic loss. CityAI requires metrics that connect micro behaviour to macro urban performance. Scaling heterogeneous agents alone, however, does not generate CityAI. The next challenge is understanding how interactions among these agents give rise to collective urban intelligence.

4.2 From multi-agent coordination to collective urban intelligence

The second challenge is to move from coordination among agents to collective urban intelligence. Existing multi-agent systems can divide tasks, exchange information, negotiate objectives, and coordinate local actions [16,18,38]. These capacities are useful, but they do not necessarily produce intelligence at the scale of the city.

A system may remain fragmented even when its components are individually capable. A transport agent may reduce congestion, an energy agent may balance demand, and an emergency agent may allocate resources, while the interactions among these actions generate new risks. Local optimization may transfer congestion to another district, increase energy use, deepen inequality, or reduce resilience. CityAI must therefore recognize cross-domain interdependence rather than merely improve isolated subsystems.

Collective urban intelligence emerges through relationships among heterogeneous actors with different knowledge, objectives, timescales, and authority. Residents, planners, public institutions, infrastructure systems, businesses, robots, and artificial agents do not perceive or influence the city in the same way. Some actors may observe conditions, some may recommend actions, some may allocate resources, and others may possess legitimate authority to intervene.

Future research must therefore examine how agent roles are defined, how conflicts are represented, and how decisions are coordinated without collapsing urban plurality into a single optimization objective. Urban intelligence should not be measured by the elimination of disagreement [14,17,18]. Conflict may reveal legitimate differences in values, knowledge, and interests. The objective is to increase the capacity of urban systems to recognize interdependence, expose trade-offs, negotiate among perspectives, and coordinate appropriate responses.

Collective intelligence must also include existing forms of human and institutional intelligence. Communities already coordinate through social networks, professional knowledge, informal norms, markets, civic organizations, and political institutions. Artificial agents should not replace these arrangements. Their role should be to improve the visibility of interactions, explore alternative outcomes, support deliberation, and reveal consequences that no single actor can fully observe.

The key shift is therefore from asking whether agents can collaborate to asking whether their relationships can produce coherent urban capacities. CityAI depends less on the intelligence of the strongest agent than on the quality, legitimacy, and adaptability of relationships among many limited intelligences. In this sense, collective urban intelligence should be understood as an emergent property arising from interactions among heterogeneous agents rather than from the capability of any individual agent.

4.3 From urban prediction to urban reasoning

The third challenge is the transition from prediction to urban reasoning. Urban AI has achieved substantial progress in forecasting traffic, energy use, environmental risk, mobility, land-use change, and infrastructure failure [11,14]. Prediction estimates what is likely to happen. CityAI must additionally ask why conditions are emerging, what actions are possible, who will be affected, which constraints apply, and how interventions should be justified.

Urban reasoning is inherently cross-systemic. A flood-risk prediction does not determine an appropriate response. A CityAI system must relate hazard estimates to drainage capacity, road accessibility, vulnerable populations, emergency resources, hospital capacity, public communication, legal responsibility, and expected human behaviour. Different agents may hold partial knowledge and conflicting priorities. The task is therefore not to generate one answer, but to construct an auditable process connecting evidence, rules, models, values, and authority.

Four forms of reasoning are particularly important. Causal reasoning is required to identify mechanisms, evaluate interventions, and construct counterfactuals [4]. Spatial reasoning is necessary because urban decisions depend on distance, topology, networks, morphology, scale, and spatial heterogeneity. Normative reasoning is needed because safety, efficiency, equity, development, ecological protection, and cultural continuity may conflict. Rule-grounded reasoning is essential because urban actions are constrained by laws, plans, standards, procedures, and institutional permissions.

Reasoning-based CityAI will therefore be multi-model and multi-agent. A planning or governance agent may use GIS, transport models, environmental simulations, regulatory databases, and stakeholder inputs. One agent may generate an alternative, another verify compliance, another evaluate distributional impacts, and a human authority may resolve unresolved value conflicts.

The quality of CityAI should not be assessed through fluent outputs alone. Urban reasoning must be traceable, spatially grounded, evidence-based, uncertainty-aware, and open to challenge. This also changes professional practice. Planners and urban managers will increasingly need to design reasoning workflows, define agent roles, inspect evidence chains, and determine which decisions must remain under human and institutional control. Such reasoning ultimately requires continuous interaction with the physical city, motivating the transition from Digital Twins toward living urban models.

4.4 From Digital Twins to living urban models

The fourth challenge is the transformation of Digital Twins into living urban models. Urban Digital Twins increasingly combine three-dimensional representation, real-time sensing, simulation, and prediction [20–25,40,43,44,52,60]. These capabilities provide an important foundation for CityAI, but representation alone does not create intelligence.

A living urban model should be understood as a continuously evolving and explicitly uncertain hypothesis about how the city works. It must connect observation, simulation, intervention, evaluation, and revision. Rather than presenting one visually precise representation, it should reveal where knowledge is incomplete, where models disagree, and which assumptions remain uncertain.

Living urban models must integrate multiple spatial and temporal scales. Seconds may matter for traffic operations, hours for heat exposure, days for emergency response, and years for land-use and infrastructure change [57]. They must also incorporate human and institutional behaviour. A flood model may estimate water depth, but evacuation outcomes depend on trust, communication, household structure, mobility resources, social networks, and government action.

CityAI also requires interaction among multiple specialized models rather than one supposedly complete model of the city. Transport, climate, energy, social, economic, and institutional models use different assumptions and resolutions [30,31]. Their outputs must be connected without erasing uncertainty or forcing artificial consistency.

Most importantly, living urban models must close the loop with the physical and institutional city. Simulations should generate testable propositions; bounded interventions should produce observable outcomes; and differences between expected and observed results should update models and strategies. This creates an iterative methodology of simulate, deliberate, intervene, observe, compare, and revise.

Such experimentation must remain proportionate and governable. Cities cannot be treated as unrestricted laboratories. High-risk interventions should remain in computational environments, while real-world tests should be reversible, transparent, and institutionally authorized. A living urban model becomes part of CityAI only when its learning loop is connected to legitimate urban decision-making. As these models increasingly influence human behaviour, the relationship between people and AI also becomes evolutionary rather than supervisory.

4.5 From human-in-the-loop to human–AI co-evolution

The fifth challenge is to move beyond a static understanding of human oversight. Human-in-the-loop approaches generally assume that an AI system performs a task and a person reviews or corrects its output [18,38,50]. Urban systems are more recursive. People do not merely supervise AI; they also adapt to it, and AI systems subsequently learn from environments already altered by previous interventions.

Navigation systems reshape travel behaviour and neighbourhood traffic. Recommendation systems affect destination popularity. Automated enforcement changes public behaviour. Generative planning alters professional workflows. Robotic services transform expectations of public space. AI-supported policies may influence investment, development, and institutional organization.

These interactions create co-evolution. Humans change AI systems; AI systems change human behaviour; both reshape urban environments and institutions; and those changes influence subsequent human and artificial decisions [9,45]. A system that performs well initially may generate long-term effects that undermine its purpose. Route guidance may transfer congestion to residential streets. Cooling interventions may attract more users and increase local demand. Digital services may improve access for some groups while reducing human support for others.

Co-evolution also occurs within institutions. Governments may reorganize workflows around AI, create new professional roles, change procurement practices, and revise standards or regulations. These processes influence which technologies become embedded in urban infrastructure and whose knowledge they privilege.

Future CityAI research therefore requires longitudinal evaluation. Short-term demonstrations cannot reveal dependency, deskilling, behavioural adaptation, institutional lock-in, trust erosion, or emerging inequality [10,47,48,55]. Researchers must examine how urban actors change over months and years, and whether the benefits and burdens of AI mediation are distributed fairly.

Human agency must remain central. Co-evolution should not mean passive social adaptation to technology. Residents, professionals, and institutions must be able to shape the objectives, boundaries, and permissions of CityAI. Participation should begin before systems are deployed, including decisions about what should be sensed, which values should guide evaluation, what may be automated, and which decisions must remain open to public and professional judgement. Ultimately, sustained human–AI co-evolution raises the broader question of when an AI-enabled city can truly be regarded as CityAI.

4.6 From the AI City to CityAI

The previous five directions describe complementary research frontiers. Taken together, they provide the basis for asking a more integrative question: under what conditions can an AI-enabled city be regarded as CityAI? The Smart City can be understood as a city equipped with digital technologies and data infrastructures [19]. The AI City advances this paradigm by embedding Artificial Intelligence into urban perception, judgement, coordination, and operation. CityAI represents a further shift: the formation of urban-scale intelligence through continuous interaction among heterogeneous human and artificial agents, infrastructures, institutions, and environments.

This transition is not a rejection of the AI City. It extends its central proposition under the conditions created by foundation models, agentic AI, multi-agent systems, Digital Twins, and embodied intelligence. The AI City asks how Artificial Intelligence can become a structural capacity of the city. CityAI asks when that capacity becomes distributed, relational, adaptive, governable, and emergent.

Not every AI-enabled platform, digital twin, or multi-agent system should therefore be described as CityAI. To distinguish CityAI from these related systems, we specify five operational conditions at the urban-system level. First, CityAI must be cross-domain and relational rather than confined to one application. Second, it must connect perception, interpretation, reasoning, coordination, action, evaluation, and evolution in a closed but contestable loop. Third, it must produce system-level capacities that cannot be reduced to the performance of one agent. Fourth, its roles, permissions, and responsibilities must correspond to legitimate institutional authority. Fifth, it must be evaluated across individuals, groups, places, systems, and time. These conditions should not be interpreted as a binary certification of CityAI, but as analytical dimensions for assessing the extent to which urban intelligence becomes relational, emergent, governable, and system-wide.

These conditions imply that CityAI should be polycentric rather than dependent on a single omniscient model. Specialized agents, communities, institutions, and infrastructures should retain local knowledge and authority while participating in broader coordination. CityAI should make objectives and trade-offs explicit, distinguish prediction from normative judgement, and recognize different levels of authority between observation, recommendation, decision, and physical intervention.

CityAI should also be adaptive without becoming directionless. Strategies and models may change, while fundamental rights, public values, and legal protections require stability [46,54,55]. It should preserve the possibility of contestation because disagreement is an intrinsic condition of urban life [15]. The purpose of CityAI is not to eliminate conflict, but to increase the city’s capacity to perceive, explain, deliberate over, and coordinate through complexity.

The resulting research programme requires heterogeneous agent architectures, cross-scale evaluation frameworks, urban reasoning tools, living urban models, governance protocols, and longitudinal real-world testbeds. Campuses, public spaces, districts, and selected urban services may provide bounded settings for testing multi-agent coordination, institutional decision-making, physical response, and learning without prematurely claiming citywide intelligence.

The historical trajectory of urban computation has moved from calculating the city, to simulating it, sensing it, learning from it, generating alternatives for it, and connecting models to goals and actions [18]. CityAI asks a further question: can an urban system develop the capacity to understand, coordinate, evaluate, and revise itself through the interaction of its human and artificial inhabitants?

The answer is cautiously affirmative, but conditional. Technical capability alone will not create CityAI [27–29]. Human needs, planning judgement, operational governance, complexity efficiency, adaptive learning, and hybrid intelligence must converge [16]. The next frontier of artificial intelligence may therefore not lie in building larger models or more autonomous agents, but in understanding how intelligence itself becomes an emergent property of the city—a paradigm referred to here as CityAI.

References

[1]

Wiener N. Cybernetics: or control and communication in the animal and the machine. Cambridge: MIT Press, 1948

[2]

Simon H A. The sciences of the artificial. Cambridge: MIT Press, 1969

[3]

Forrester J W. Urban dynamics. Cambridge: MIT Press, 1969

[4]

Zheng Y , Capra L , Wolfson O . et al. Urban computing: concepts, methodologies, and applications. ACM Transactions on Intelligent Systems and Technology, 2014, 5(3): 38

[5]

Kitchin R . The real-time city? Big data and smart urbanism. GeoJournal, 2014, 79(1): 1–14

[6]

Caragliu A , Del Bo C , Nijkamp P . Smart cities in Europe. Journal of Urban Technology, 2011, 18(2): 65–82

[7]

Hollands R G. . Will the real smart city please stand up?. City, 2008, 12(3): 303–320

[8]

Shelton T , Zook M , Wiig A . The ‘actually existing smart city’. Cambridge Journal of Regions, Economy and Society, 2015, 8(1): 13–25

[9]

Allam Z , Dhunny Z A . On big data, artificial intelligence and smart cities. Cities, 2019, 89: 80–91

[10]

Bibri S E , Krogstie J . Smart sustainable cities of the future: an extensive interdisciplinary literature review. Sustainable Cities and Society, 2017, 31: 183–212

[11]

LeCun Y , Bengio Y , Hinton G . Deep learning. Nature, 2015, 521(7553): 436–444

[12]

Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach, USA, 2017, 6000–6010

[13]

Brown T B, Mann B, Ryder N, et al. Language models are few-shot learners. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. Vancouver, Canada, 2020, 159

[14]

Park J S, O'Brien J, Cai C J, et al. Generative agents: interactive simulacra of human behavior. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. San Francisco, USA, 2023, 2. doi: 10.1145/3586183.3606763

[15]

Cugurullo F . Urban artificial intelligence: from automation to autonomy in the smart city. Frontiers in Sustainable Cities, 2020, 2: 38

[16]

Bonabeau E . Agent-based modeling: methods and techniques for simulating human systems. Proceedings of the National Academy of Sciences of the United States of America, 2002, 99(S3): 7280–7287

[17]

Grimm V , Berger U , Bastiansen F . et al. A standard protocol for describing individual-based and agent-based models. Ecological Modelling, 2006, 198(1−2): 115–126

[18]

Dellermann D , Ebel P , Söllner M . et al. Hybrid intelligence. Business & Information Systems Engineering, 2019, 61(5): 637–643

[19]

Wu S Z. The AI city. Singapore: Springer, 2025

[20]

Batty M . Digital twins. Environment and Planning B: Urban Analytics and City Science, 2018, 45(5): 817–820

[21]

Dembski F , Wössner U , Letzgus M . et al. Urban digital twins for smart cities and citizens: the case study of Herrenberg, Germany. Sustainability, 2020, 12(6): 2307

[22]

Shahat E , Hyun C T , Yeom C . City digital twin potentials: a review and research agenda. Sustainability, 2021, 13(6): 3386

[23]

Weil C , Bibri S E , Longchamp R . et al. Urban digital twin challenges: a systematic review and perspectives for sustainable smart cities. Sustainable Cities and Society, 2023, 99: 104862

[24]

Lehtola V V , Koeva M , Elberink S O . et al. Digital twin of a city: review of technology serving city needs. International Journal of Applied Earth Observation and Geoinformation, 2022, 114: 102915

[25]

Xia H S , Liu Z S , Efremochkina M . et al. Study on city digital twin technologies for sustainable smart city design: a review and bibliometric analysis of geographic information system and building information modeling integration. Sustainable Cities and Society, 2022, 84: 104009

[26]

Schelling T C . Dynamic models of segregation. The Journal of Mathematical Sociology, 1971, 1(2): 143–186

[27]

Batty M. Cities and complexity: understanding cities with cellular automata, agent-based models, and fractals. Cambridge: MIT Press, 2005

[28]

Holland J H. Hidden order: how adaptation builds complexity. Reading: Addison-Wesley, 1995

[29]

Epstein J M, Axtell R. Growing artificial societies: social science from the bottom up. Washington: Brookings Institution Press, 1996

[30]

Crooks A , Castle C , Batty M . Key challenges in agent-based modelling for geo-spatial simulation. Computers, Environment and Urban Systems, 2008, 32(6): 417–430

[31]

An L . Modeling human decisions in coupled human and natural systems: review of agent-based models. Ecological Modelling, 2012, 229: 25–36

[32]

Torrens P M , Benenson I . Geographic automata systems. International Journal of Geographical Information Science, 2005, 19(4): 385–412

[33]

O'Sullivan D , Haklay M. . Agent-based models and individualism: is the world agent-based?. Environment and Planning A: Economy and Space, 2000, 32(8): 1409–1425

[34]

Mitchell W J. City of bits: space, place, and the infobahn. Cambridge: MIT Press, 1995

[35]

Townsend A M. Smart cities: big data, civic hackers, and the quest for a new Utopia. New York: W. W. Norton, 2013

[36]

Goodchild M F . Citizens as sensors: the world of volunteered geography. GeoJournal, 2007, 69(4): 211–221

[37]

Vanolo A . Smartmentality: the smart city as disciplinary strategy. Urban Studies, 2014, 51(5): 883–898

[38]

Seeber I , Bittner E , Briggs R O . et al. Machines as teammates: a research agenda on AI in team collaboration. Information & Management, 2020, 57(2): 103174

[39]

Zheng Y , Lin Y M , Zhao L . et al. Spatial planning of urban communities via deep reinforcement learning. Nature Computational Science, 2023, 3(9): 748–762

[40]

Ketzler B , Naserentin V , Latino F . et al. Digital twins for cities: a state of the art review. Built Environment, 2020, 46(4): 547–573

[41]

Ritter H , Herzog O , Rothermel K . et al. City models: past, present and future prospects. Frontiers of Urban and Rural Planning, 2025, 3(1): 7

[42]

Batty M , Axhausen K W , Giannotti F . et al. Smart cities of the future. The European Physical Journal Special Topics, 2012, 214(1): 481–518

[43]

Supianto A A , Nasar W , Margrethe Aspen D . et al. An urban digital twin framework for reference and planning. IEEE Access, 2024, 12: 152444–152465

[44]

Li D R , Yu W B , Shao Z F . Smart city based on digital twins. Computational Urban Science, 2021, 1(1): 4

[45]

Cardullo P , Kitchin R . Being a ‘citizen’ in the smart city: up and down the scaffold of smart citizen participation in Dublin, Ireland. GeoJournal, 2019, 84(1): 1–13

[46]

UN-Habitat. AI & Cities: Risks, Applications and Governance. Nairobi: United Nations Human Settlements Programme, 2022

[47]

Yigitcanlar T , Cugurullo F . The sustainability of artificial intelligence: an urbanistic viewpoint from the lens of smart and sustainable cities. Sustainability, 2020, 12(20): 8548

[48]

Kitchin R , Dodge M . The (In)security of smart cities: vulnerabilities, risks, mitigation, and prevention. Journal of Urban Technology, 2019, 26(2): 47–65

[49]

Saarloos D , Arentze T , Borgers A . et al. A multiagent model for alternative plan generation. Environment and Planning B: Planning and Design, 2005, 32(4): 505–522

[50]

Jarrahi M H . Artificial intelligence and the future of work: human-AI symbiosis in organizational decision making. Business Horizons, 2018, 61(4): 577–586

[51]

Waddell P . UrbanSim: modeling urban development for land use, transportation, and environmental planning. Journal of the American Planning Association, 2002, 68(3): 297–314

[52]

Nochta T , Wan L , Schooling J M . et al. A socio-technical perspective on urban analytics: the case of city-scale digital twins. Journal of Urban Technology, 2021, 28(1−2): 263–287

[53]

Wu Z Q , Pan Y H , Ye Q M . et al. The City Intelligence Quotient (City IQ) evaluation system: conception and evaluation. Engineering, 2016, 2(2): 196–211

[54]

OECD. Artificial Intelligence for Advancing Smart Cities. Paris: OECD, 2025

[55]

Meijer A , Bolívar M P R . Governing the smart city: a review of the literature on smart urban governance. International Review of Administrative Sciences, 2016, 82(2): 392–408

[56]

Wang H H , Shi W Y , He W L . et al. Simulation of urban transport carbon dioxide emission reduction environment economic policy in China: an integrated approach using agent-based modelling and system dynamics. Journal of Cleaner Production, 2023, 392: 136221

[57]

Herzog R H, Degkwitz T, Verma T. The urban model platform: a public backbone for modeling and simulation in urban Digital Twins. arXiv preprint: arXiv: 2506.10964, 2025

[58]

Ratti C, Claudel M. The City of tomorrow: sensors, networks, hackers, and the future of urban life. New Haven: Yale University Press, 2016

[59]

Couclelis H . The construction of the digital city. Environment and Planning B: Planning and Design, 2004, 31(1): 5–19

[60]

Batty M . Digital Twins in city planning. Nature Computational Science, 2024, 4(3): 192–199

[61]

Batty M. The computable city: histories, technologies, stories, predictions. Cambridge: MIT Press, 2024

[62]

Zhang W J, Han J D, Xu Z, et al. Urban foundation models: a survey. In: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. Barcelona, Spain, 2024, 6633–6643

[63]

Jacobs J. The death and life of great American cities. New York: Random House, 1961

[64]

Gehl J. Cities for people. Washington: Island Press, 2010

[65]

UN-Habitat. Global Assessment of Responsible AI in Cities. Nairobi: United Nations Human Settlements Programme, 2024

[66]

Yang S, Li J, Biljecki F. Reasoning is all you need for urban planning AI. arXiv preprint: arXiv: 2511.05375, 2025

[67]

Qian K J, Mao L J, Liang X, et al. AI agent as urban planner: steering stakeholder dynamics in urban planning via consensus-based multi-agent reinforcement learning. arXiv preprint: arXiv: 2310.16772, 2023

[68]

Han J B, Ning Y S, Yuan Z R, et al. Large language model powered intelligent urban agents: concepts, capabilities, and applications. arXiv preprint: arXiv: 2507.00914, 2025

[69]

Tiwari A . Conceptualising the emergence of Agentic Urban AI: from automation to agency. Urban Informatics, 2025, 4: 13

Rights & permissions

The Author(s) 2026. This article is published by Higher Education Press.

PDF (4881KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/

〈 〉