Optimising LoRA for fine-tuning lightweight LLM expert agents in cross-regional collaboration in the construction industry: Evidence from the Greater Bay Area

Liqun XIANG , Yaxin CAO , Geoffrey Qiping SHEN , Binwei GAO

Eng. Manag ›› 2026, Vol. 13 ›› Issue (3) : 650 -672.

PDF (12967KB)
Eng. Manag ›› 2026, Vol. 13 ›› Issue (3) :650 -672. DOI: 10.1007/s42524-026-5274-4
Construction Engineering and Intelligent Construction
RESEARCH ARTICLE
Optimising LoRA for fine-tuning lightweight LLM expert agents in cross-regional collaboration in the construction industry: Evidence from the Greater Bay Area
Author information +
History +
PDF (12967KB)

Abstract

Under the “One Country, Two Systems, Three Legal Jurisdictions” framework, cross-regional collaboration in the construction industry of the Guangdong–Hong Kong–Macao Greater Bay Area (GBA) is more complex than other urban agglomerations, and this makes traditional methods relying on manual analysis of influencing factors inefficient and subjective. While existing large language models (LLMs) can meet the needs of intelligent applications, they lack specific domain knowledge. Considering the intelligent advantages of LLMs, this research proposes a lightweight expert agent through providing an optimised low-rank adaptation (LoRA) model SVDSR-LoRA, integrating domain knowledge using the fine-tuning method. Experiments on the 1.5b lightweight base model of qwen2.5 and deepseek-r1 show that the proposed SVDSR-LoRA training method can increase the mid-term convergence speed by 36%-50% compared with the standard LoRA method. A lightweight multi-expert agent influencing factor analysis system is constructed to simulate a collaborative analysis environment with experiments showing the hit rate in high-frequency influencing factors reached 100%. The comprehensive evaluation scored at 4.03/5.0 and demonstrated the capability in revealing the external environmental factors influencing cross-regional collaboration in the construction industry of the GBA systematically, which was significantly better than that of a single model (2.52-3.22/5.0). The proposed SVDSR-LoRA training model and the MAS system establishing pattern can provide a reference for rapid and effective lightweight agent system construction methods for intelligent analysis tasks of cross-regional collaboration in the construction industry of the GBA, as well as other similar LLM-based cross-regional collaboration research

Graphical abstract

Keywords

Guangdong–Hong Kong–Macao Greater Bay Area (GBA) / cross-regional collaboration / construction industry / large language models (LLM) / lightweight expert agent

Cite this article

Download citation ▾
Liqun XIANG, Yaxin CAO, Geoffrey Qiping SHEN, Binwei GAO. Optimising LoRA for fine-tuning lightweight LLM expert agents in cross-regional collaboration in the construction industry: Evidence from the Greater Bay Area. Eng. Manag, 2026, 13 (3) : 650-672 DOI:10.1007/s42524-026-5274-4

登录浏览全文

4963

注册一个新账户 忘记密码

1 Introduction

An urban agglomeration refers to an economic and social spatial entity composed of several closely connected and functionally complementary cities and their surrounding areas, serving as an important carrier for regional economic development. The Chinese government places great emphasis on urban agglomerations as a key approach to promoting regional coordinated development and advancing new-type urbanisation, and has issued policies including “National New-type Urbanisation Plan (2014–2020)” and the “Guidelines on Cultivating and Developing Modern Metropolitan Areas.” The GBA, encompassing Hong Kong, Macao, and nine cities in Guangdong Province, is one of the most open and economically dynamic regions in China and serves as a key engine for national development. On July 1, 2017, the National Development and Reform Commission, together with the governments of Hong Kong, Macao, and Guangdong, jointly signed the “Framework Agreement on Deepening Guangdong–Hong Kong–Macao Cooperation and Promoting the Development of the Greater Bay Area” in Hong Kong. On February 18, 2019, the Central Committee of the Communist Party of China and the State Council issued the “Outline Development Plan for the Guangdong–Hong Kong–Macao Greater Bay Area.” Under the leadership of the Chinese government, cities within the GBA have made substantial and effective progress in both hard and soft connectivity, including infrastructure integration, alignment of institutional frameworks and standards, as well as cultural and interpersonal exchanges. As a result, economic and social linkages have become increasingly close, and the overall framework for the GBA has begun to be formed (Wu, 2023). As a mega-urban agglomeration under China’s basic state policy of “One Country, Two Systems,” the GBA encompasses three distinct customs territories, legal systems, and currencies. GBA’s cross-regional collaboration mechanisms are unprecedented in the international context. This unique institutional configuration presents both challenges to, and distinctive advantages for, the region’s international development (Xu and Wu, 2019). In 2023, the GBA’s total economic output exceeded 14 trillion RMB, contributing one-ninth of the national GDP while accounting for less than 0.6% of the country’s land area (Ye and Wu, 2024).

As a pillar industry of the national economy and a fundamental driver of regional development, the construction industry plays an irreplaceable role in promoting the high-quality development of the GBA (Alaghbari, Al-Sakkaf and Sultan, 2019; Raza et al., 2021). Within the policy framework of building a “unified national market,” the construction sector has emerged as a crucial domain for dismantling local protectionism. A growing number of construction enterprises have entered the GBA, contributing to its infrastructural and economic development. In recent years, the proportion of cross-regional collaboration in construction projects within the GBA has increased significantly. Through interregional collaboration in the construction industry, the GBA has achieved a more efficient resource integration and a more optimised industrial chain layout, thereby enhancing the region’s overall competitiveness (Pan, 2024). However, cross-regional collaboration in the construction industry of the GBA still faces urgent institutional barriers, including significant discrepancies in technical standards, limited mutual recognition of construction qualifications, and numerous restrictions on cross-regional bidding. These challenges have led to a range of practical difficulties.

There are significant differences among Hong Kong, Macao, and China’s mainland in terms of legal systems, construction procedures, and qualification requirements. These divergences not only increase the complexity of cross-regional collaboration but also impose greater demands on policy coordination and institutional innovation (Zhang, 2018). Cross-regional collaboration in the construction industry relies on the complex coupling of institutional, market, and technological systems, and is influenced by external environmental factors such as the degree of policy alignment, information asymmetry, and disparities in standardisation and rating systems (Xie et al., 2022). However, existing studies mainly adopt qualitative research methods such as surveys, expert interviews, and case studies. While these approaches are effective in identifying key factors affecting cross-regional collaboration and provide a certain degree of empirical support, the research process typically involves repeated visits to multiple locations and the collection of large amounts of information, which is time-consuming and labor-intensive (Varshney et al., 2016). As a result, such methods struggle to quantitatively assess the interaction effects among influencing factors or to capture the dynamic interplay between policy iterations and stakeholder behavior (Wu et al., 2021). Besides, the related studies were conducted by different research teams, each accessing multiple field sites across various cities and regions. Without cross-checking—and because questionnaires are often broad or lack specificity—respondents may provide similar answers, causing different teams to gather overlapping data. As a result, the conclusions across these studies repeatedly identify similar influencing factors of cross-regional collaboration, requiring substantial manual effort to extract and filter meaningful information. As for individual interviews, respondents may give lengthy, wide-ranging answers to pre-set questions that are not only imprecise but also prone to going off-topic, leaving only a small portion of each answer relevant to the specific question. This further reduces the amount of useful content and again necessitates manual extraction. Furthermore, China has a vast engineering and construction market that generates an enormous volume of data, yet the amount stored is only about 7% of that in North America. Of the limited data that is preserved, most exists as scattered files dispersed across archives and hard drives, resulting in a data utilization rate of less than 0.4% (Ding, 2020). When addressing the complex challenges of the cross-regional collaboration in the construction industry of the GBA, traditional manual approaches struggle to efficiently identify critical variables or optimise collaborative mechanisms. Therefore, it is necessary to explore more efficient and intelligent system identification and prediction methods.

In recent years, with the rapid development of deep learning technology and its successful application in tasks such as speech recognition, machine translation, and sentiment analysis, natural language processing (NLP) technology has made significant progress and played an important role in multiple fields. As a revolutionary product brought about by NLP technology, large language models (LLMs) can achieve zero-shot or few-shot feature learning and training, and possess powerful language understanding, reasoning, and generation capabilities. They have not only significantly improved the performance of tasks such as machine translation, sentiment analysis, and question-answering systems (Hao et al., 2024). For example, Kasneci et al. (2023) demonstrated the great potential of ChatGPT in personalised learning, automatic assessment, and intelligent tutoring in the field of education; Pavlik (2023) explored how generative AI can revolutionise the workflow of news production, fact-checking, and content creation in the field of journalism; and Mahowald et al. (2024) systematically analyzed the performance differences between LLMs in terms of formal and functional language capabilities in the field of linguistics.

However, unlike fields such as education, which have abundant, well-structured, and easily extractable texts, construction industry-related texts are less structured, and existing LLMs still lack sufficient vertical domain knowledge in this area. As a result, corresponding distilled models perform poorly on construction-related tasks. Furthermore, most LLM-based applications rely on a single LLM directly, which leads to suboptimal results when handling complex tasks involving multiple subtasks. Given the feasibility of LLMs in general cognition and reasoning, this research makes three principal theoretical contributions to the field of cross-regional collaboration in the construction industry. First, a SVDSR-LoRA model is proposed, which incorporates a coupled module integrating singular value decomposition (SVD) initialisation and a squared rank (SR) scaling factor into the conventional LoRA training framework. This design provides a theoretical basis for accelerating mid-term convergence and enhancing training stability of lightweight LLMs, addressing the inherent uncertainty of random noise initialisation and the gradient vanishing problem associated with traditional LoRA scaling factors. Second, a methodological framework is established for constructing lightweight domain-specific expert agents by fine-tuning distilled LLMs with vertical domain knowledge. This approach bridges the gap between general-purpose LLMs and the specialized knowledge requirements of the construction industry, offering a transferable paradigm for domain adaptation of lightweight models in knowledge-intensive fields. Third, a lightweight multi-agent system (MAS) architecture that leverages PESTEL-framed expert agents is developed to simulate multi-domain collaborative analysis. This architecture provides a systematic framework for decomposing complex cross-regional collaboration analysis into structured multi-dimensional expert discussions enhanced by PESTEL theoretical framework, offering a methodological reference for intelligent factor identification in cross-regional collaboration research. In summary, this research advances understanding of how AI-driven approaches can be applied to analyze and optimise cross-regional collaboration mechanisms in the construction industry of the GBA and comparable contexts.

2 Literature review

2.1 The Influencing factors of the cross-regional collaboration

Cross-regional collaboration refers to the governance process in which multiple stakeholders from adjacent regions integrate resources and cooperate through formal or informal institutional arrangements based on shared interests, aiming to achieve mutual economic and social development (Fang, 2014). From the perspective of intergovernmental relations, Wright (1997) classified intergovernmental policies into four types: boundary or jurisdictional, developmental and distributive, regulatory, and redistributive. Building on this theoretical paradigm, Chinese scholars have proposed corresponding frameworks such as boundary agreements, development agreements, regulatory agreements, and redistributive agreements. These have evolved into a locally adaptive governance pathway system, constituting the core institutional tools for cross-regional collaboration in China (Su, 2015).

Existing research has examined factors influencing cross-regional collaboration from multiple dimensions, offering important theoretical foundations for optimising collaborative mechanisms. In terms of policy and regulatory factors, local governments often adopt protectionist strategies to safeguard regional interests, and current administrative structures create jurisdictional barriers that hinder cross-regional collaboration. In this context, Ma et al. (2023) explored the application of cross-regional legislation in three key areas: natural resource conservation and utilization, cross-regional development and construction, and spatial governance; The research highlights the practical significance of cross-regional legislation in land-use planning and proposes pathways to improve legislative procedures. From the perspective of socioeconomic factors, population agglomeration strengthens both consumer and labor market and enhances social proximity, thereby supporting the high-quality development of cross-regional collaboration (Zhang et al., 2023). However, the uneven spatial distribution of information, human resources, and other critical inputs leads to misalignment in governance capacities among local governments, impeding effective intergovernmental collaboration (Hu, 2022). While capital accumulation often serves as a key driver of cross-regional collaboration, imbalances in benefit distribution remain a fundamental source of conflict (Zhang et al., 2023). In response, Xiang and Pang (2021) advocate for a triadic coordination mechanism integrating “actors, institutions, and interests” to enhance the efficiency of cross-regional collaboration. Regarding environmental and technological factors, natural geographic barriers may foster “warlord-style” economies that hinder the integrated development of economic belts (Yuan et al., 2018). Technological linkages between industries can facilitate collaborative innovation, yet administrative boundaries pose significant barriers to knowledge spillover, thereby constraining the deepening of cross-regional collaboration (Su et al., 2021; Tang et al., 2022).

Beyond theoretical explorations of influencing factors, empirical cases at various scales further reveal the combined effects of these variables. For instance, the Yangtze River Delta region features a large economic scale and diverse industrial base. The central government’s support for integrated connectivity has fostered deep intra-regional collaboration, yet limitations in knowledge flow have constrained the diversity of regional development (Su et al., 2021). In contrast, the GBA displays a polycentric collaboration network, with collaboration focusing on economic, social, and institutional dimensions. The integration of urban, demographic, capital, and institutional factors has facilitated the region’s advancement toward integrated development (Zhang et al., 2022; Zhang et al., 2023). In Eurasian cross-regional collaboration projects, the influence of historical and cultural contexts, geographic and transportation conditions, international dynamics, and external interference necessitates enhanced trust and communication among stakeholders to ensure sustained collaboration (Li, 2018).

While current studies have explored governance models of cross-regional collaboration, identified influencing factors such as policy, regulatory, economic, social, environmental, and technological conditions, and examined multi-scalar case practices, existing literature remains primarily focused on institutional and legislative frameworks and the role of geographical and economic elements. Although interdisciplinary approaches and case-based studies are emphasized, there is still a lack of comprehensive theoretical frameworks that integrate institutional, economic, and spatial dimensions, as well as targeted solutions addressing region-specific barriers.

2.2 Applications of NLP technology in construction project collaboration

With the continued advancement of the construction industry, the volume of textual data, including project-related standards, contract clauses, and technical reports, has grown exponentially. Traditional information processing strategies are increasingly inadequate in meeting the demands for resource sharing and information exchange during collaborative construction projects (Zhang and El-Gohary, 2016; Xu et al., 2022). NLP technologies are effective in extracting and analyzing information from unstructured text, thereby enhancing stakeholders' access to critical information (Nedeljković and Kovačević, 2017). Moreover, by leveraging similar case data, NLP can accurately identify key issues and potential opportunities in collaboration, thus offering data-driven support for optimising cooperation models and managing project risks (Zou et al., 2017; Hassan et al., 2021). As a revolutionary product brought by NLP technology, LLM is pre-trained based on large-scale text data, learns rich language knowledge and semantic information, and can generate coherent and logically reasonable text content. It has demonstrated strong text generation and comprehension capabilities in downstream tasks related to natural language processing (Qin et al., 2025).

The advancement of construction projects requires the cooperation of multiple stakeholders. Simple models are difficult to describe the complex problems faced by multiple stakeholders. Therefore, existing LLMs are usually further developed into intelligent agents to better complete special tasks (Yin and Carenini, 2025; Yue, 2025). However, as the vertical domain characteristics of the task increase, the ability of a single agent to understand the vertical domain task decreases (Li et al., 2025). Against this backdrop, MAS, which simulate continuous and dynamic interactions between agents and their environments, have garnered growing attention in the construction domain (Xiang et al., 2022). Within MAS architectures, each agent can independently process information and solve problems without relying on centralised control or synchronous computation (Lv et al., 2021). Compared to single-agent systems, MAS are capable of decomposing complex tasks through distributed intelligence and collaborative mechanisms, allocating subtasks to different agents for parallel execution and thereby significantly improving overall efficiency in multi-task scenarios (Sun et al., 2024; Yu and Dong, 2025). Traditional MAS models are mostly based on numerical simulations and use software such as AnyLogic or NetLogo, which rely on mathematical formulas rather than consciousness-based reasoning. These models often suffer from limited flexibility. Considering LLMs’ cognitive capabilities, integrating which into MAS cannot only address the above limitations but also overcome the inherent constraints of single-agent systems. This integration enables knowledge sharing and natural language communication among agents, enhances their coordination in complex tasks, and improves the decision-making efficiency of LLMs in multi-subtask scenarios (Sun et al., 2024; Yu and Dong, 2025).

However, most existing LLMs are trained on general-domain data. While high-parameter models can perform relatively simple tasks in vertical domains to some extent, the need for data protection often necessitates local deployment of models in government and enterprise settings to meet the requirements of private-domain tasks. However, the computational demands of full-parameter LLMs are typically incompatible with such specialized domain applications. On the other hand, distilled low-parameter models, due to their performance limitations, also struggle to adequately address the needs of complex domains like the construction industry. Therefore, in the context of cross-regional collaboration within the GBA’s construction sector, it is essential to inject task-specific knowledge to enhance the capabilities of lightweight models. Given the importance of intelligent factor identification for analyzing cross-regional collaboration and the advantages of LLMs in cognitive flexibility and adaptability, the authors first propose the SVDSR-LoRA model, based on existing LoRA training frameworks, to optimise model training efficiency. Building on the SVDSR-LoRA model, the authors fine-tune the Qwen2.5:1.5B model using domain-specific knowledge related to cross-regional collaboration in the construction industry in the GBA, thereby creating a lightweight expert agent. Furthermore, a lightweight MAS is constructed based on these expert agents to intelligently analyze and uncover key influencing factors in such collaborations. This approach aims to offer new insights and references for AI-driven studies of cross-regional collaboration within the construction industry.

3 Methodology

When building an intelligent analysis system to match the relevant specific tasks in the cross-regional collaboration in the construction industry of the GBA, vertical field tasks such as the construction industry often have a higher reliance on professional knowledge. In the context of collaboration in the GBA with highly heterogeneous laws and systems, the model needs to have a deeper understanding and application of field-specific knowledge. The authors take the intelligent identification of factors affecting cross-regional collaboration in the construction industry in the GBA as a case study, and explores how to optimise the existing model training scheme while ensuring the lightweight characteristics of the model. The SVDSR-LoRA solution is proposed with a coupled module through integrating both singular value decomposition (SVD) and squared rank (SR) optimising module, then the relevant field knowledge is used to enhance the adaptability of the lightweight LLM in professional tasks. Specifically, this research consists of data collection and preprocessing, an SVD-based optimisation module, and an SR-based optimisation module, as shown in Fig. 1.

3.1 Data collection and influencing factor framework construction

As previously mentioned, existing LLMs are primarily trained on general-domain data and lack specialized knowledge in the construction industry. As a result, their lightweight counterparts cannot be directly applied to specific vertical-domain tasks such as identifying the influencing factors of cross-regional collaboration in the construction industry of the GBA (Lee and Lee, 2024). Therefore, it is necessary to enrich existing lightweight models with domain-specific textual knowledge to customise lightweight expert agents for such applications.

Given the high textual quality of policy documents, consulting reports, and academic literature—and considering that the GBA is a unique urban agglomeration under China’s “One Country, Two Systems” basic state policy, involving both Chinese and English language research—the authors drew on sources in both languages. In addition to key policy and consulting documents such as the “Outline Development Plan for the Guangdong–Hong Kong–Macao Greater Bay Area” issued by the Central Committee of the Communist Party of China and the State Council, and the “GBA ESG Action Report” released by China Media Group, the authors also collect academic texts from both Chinese and international databases. Specifically, Chinese publications were sourced from China National Knowledge Infrastructure, focusing on journals categorised as Peking University Core and CSSCI. English publications were retrieved from the Web of Science database. The search targeted literature published between 2005 and 2024 using keywords such as跨区域合作 (cross-regional/cross-border cooperation/collaboration) and 大湾区 (GBA). After screening and removing irrelevant entries, a total of 39 Chinese journal articles (from an initial pool of 584) and 15 English journal articles (from an initial pool of 340) were identified as closely related to cross-regional collaboration in the construction industry within the GBA.

In practical applications of LLMs, if conditional constraints are not imposed, variations in the model’s temperature parameter can lead to overly volatile outputs. As prior research has demonstrated the effectiveness of the PESTEL framework, the authors adopt PESTEL as a top-level constraint for the expert agents’ responses when analyzing influencing factors (Pan et al., 2019). Specifically, the expert agents are guided to frame their responses within six key dimensions: Political, Economic, Social, Technological, Environmental, and Legal.

3.2 Strategies for developing a LLM-based lightweight expert agent

Although LLMs have demonstrated strong capabilities across various NLP tasks, their training on general-domain data results in a significant lack of specialized knowledge in vertical fields. In the complex scenario of cross-regional collaboration within the construction industry of the GBA, elements such as policy and regulatory coordination, cultural differences, technical standard alignment, and resource allocation mechanisms form a multidimensional and dynamic system. This system not only demands real-time responsiveness to evolving policy developments but also requires a deep understanding of long-established collaborative paradigms. Due to the professional, complex, and highly specialized nature of knowledge in the construction industry, LLMs lacking vertical-domain expertise often struggle to accurately comprehend and process relevant information. This limitation frequently leads to factual hallucinations—generating content that appears plausible but is incorrect or fabricated (He, Chen and Dai, 2025a; He, Shen and Xie, 2025b). As a result, the practical application of LLMs in the construction sector faces significant challenges, limiting their effectiveness and reliability in key areas such as architectural design, construction management, and engineering consulting (Göpfert et al., 2024; He et al., 2024), which highlights the need of domain knowledge integration of existing LLMs.

Current methodologies for developing domain-specific LLM-based agents primarily encompass two paradigms: parameter-efficient fine-tuning and retrieval-augmented generation (RAG). RAG-based approaches retrieve relevant documents from an external knowledge base during inference and integrate them into the model’s context window, thereby enabling access to updated information without model retraining. However, as the length and volume of retrievable documents increase, particularly in scenarios involving multi-source, lengthy inputs with broad semantic spans, the retrieval recall of RAG systems tends to decline, revealing inherent limitations under such conditions. Since this research examines the influencing factors of cross-regional collaboration in the construction industry of the GBA, the required domain knowledge is relatively stable, well-defined. It is predominantly derived from policy documents, academic literature, and industry reports with low update frequency. Consequently, the necessity for real-time retrieval is limited.

Compared with RAG, the fine-tuning approach enables deeper internalisation of domain knowledge extracted from pre-processed and standardised multi-source long texts directly into model parameters. This not only circumvents the latency and computational overhead associated with retrieval pipelines, but also ensures alignment with the PESTEL analytical framework, thereby generating customised and theory-grounded outputs. Within the subsequent multi-agent system, where each agent must maintain a consistent and specialized domain perspective throughout collaborative reasoning, knowledge embedded via fine-tuning yields more customised and stable expert behavior than outputs that may fluctuate due to variable retrieval results across inference calls. This internalisation strategy therefore enhances both output fidelity and system robustness in resource-constrained deployment settings. Accordingly, the authors enhance an existing LLM by fine-tuning it with construction domain-specific content, including academic literature, policy documents, and interview transcripts, thereby enriching its knowledge and improving its applicability to real-world scenarios.

Current training approaches for LLMs include full fine-tuning, adapters, prefix-tuning, and LoRA, among others. Full fine-tuning involves updating the weights of a pre-trained model using supervised data specific to a target task, and it has demonstrated strong performance across various benchmark tasks (Brown et al., 2020). This method can significantly improve the baseline model’s task-specific capabilities. However, full fine-tuning also presents four major limitations: First, it typically requires large-scale training data sets (often exceeding 1GB), which can lead to overfitting and reduced generalisation performance (Hao et al., 2024). Second, it necessitates storing the same number of parameters as the original model, which results in high storage costs when managing multiple fine-tuned model instances (Hu et al., 2021). Third, the training process is computationally intensive, demanding substantial hardware resources. Fourth, full fine-tuning is prone to catastrophic forgetting, where learning new tasks may overwrite or degrade the model’s previously acquired knowledge representations (Kumar et al., 2022).

To address the aforementioned issues, methods such as adapters—which add small additional modules—and prefix tuning—which modifies the input prefix vectors—have been proposed (Houlsby et al., 2019; Li and Liang, 2021). However, these approaches, which only adjust partial parameters or optimise external modules, often struggle to match the model quality achieved by full fine-tuning. In response, Hu et al. introduced LoRA, a fine-tuning method designed to reduce the number of trainable parameters and computational resource requirements while maintaining the model’s expressive power and adaptability (Hu et al., 2021). LoRA offers dual advantages: in terms of storage, it only requires maintaining the pre-trained model alongside task-specific low-rank matrices, significantly improving storage efficiency and enabling rapid switching between multiple tasks; While regarding computation, only a small number of low-rank parameters are optimised without needing to compute gradients or maintain optimiser states for most model parameters, it substantially lowers hardware demands and training costs. The computational expense can be reduced to approximately one-third of that required by traditional full fine-tuning (Hu et al., 2021).

Considering LoRA can reduce the number of trainable parameters and computing resource requirements while maintaining the model’s expressiveness and adaptability, it brings the possibility of training lightweight expert agents for cross-regional collaboration in the construction industry in the GBA through specific knowledge. Therefore, the authors optimise the baseline model based on LoRA to improve the LLM’s ability to analyze the factors affecting cross-regional collaboration in the construction industry.

3.3 Fine tuning approach

In the practice of fine-tuning large-scale pre-trained models, when model size and task complexity increase, traditional full-parameter fine-tuning approaches not only incur significant computational and storage overhead, but also suffer from overfitting, slow convergence, and training instability. To address this, researchers have proposed a series of PEFT methods, among which LoRA has attracted considerable attention for its simple yet efficient low-rank matrix insertion approach.

LLMs differ from traditional machine learning and deep learning models in that they have a large number of parameters, resulting in extremely high training costs. As the number of model parameters increases, this cost increases by more than a hundredfold. Therefore, it is necessary to optimise model training methods to ensure that accuracy meets requirements while reducing convergence time and optimising model effectiveness. To this end, the authors further optimise the LoRA fine-tuning model by introducing SVD singular value decomposition and a squared rank factor module to propose the SVDSR-LoRA model. This model optimises convergence speed and training stability while maintaining the required number of trainable parameters and video memory usage.

3.3.1 Singular value decomposition (SVD) optimising module

LoRA, as a parameter-efficient fine-tuning (PEFT) method, is characterized by approximating the model parameter update ΔW through a low-rank decomposition into the product of two small matrices A and B (ΔW=AB). In this approach, matrix A is initialised with Gaussian noise, matrix B is initialised to zero, and the original model parameters W are kept frozen while only A and B are trained, thereby reducing the number of trainable parameters. This process is illustrated in the yellow and green blocks of Fig. 2.

Specifically, for the original weight matrix W∈Rm×n, LoRA defines the low-rank decomposition update as follows:

ΔW=AB,A∈Rm×r,B∈Rr×n,r≪min(m,n).

The updated weight is calculated as:

W∗=W+ΔW=W+AB.

Then the forward pass can then be computed as:

Y=XW∗=X(W+AB).

Since the above weight computation design relies on random noise initialisation, it introduces uncertainty in the early training stages, often leading to slower convergence and potential difficulty in reaching the optimal solution. As training scale increases, this slower convergence can lead to substantial costs for the corresponding tasks. To improve fine-tuning convergence speed, the authors introduce a SVD module to initialise the original weight matrices, which optimises the uncertainty situation in early training and accelerating model convergence (Meng et al., 2024). The corresponding illustration is shown in the yellow blocks of Fig. 2.

The residual matrix is constructed by initialising the original weight matrix W based on its principal components weight through SVD:

Wres=USVT,U∈Rm×k,S=diag(s),s∈R⩾0min(m,n),V∈Rn×k,k=min(m,n).

For A and B, the top r largest singular values are selected to form the principal components:

A=U[:,:r]S:r,:r12∈Rm×r,

B=S:r,:r12V[:,:r]T∈Rr×n.

The newly initialised principal component weights are constructed based on the principal components A and B:

Wres=U[:,r:]S[r:,r:]V[:,r:]T∈Rm×n,

Wpri=AB.

Finally, the forward propagation is performed by combining the new principal component weights and the residual matrix:

Y=XW∗=X(Wres+Wpri)=X(U[:,r:]S[r:,r:]V[:,r:]⊤+AB).

Overall, the difference between the proposed module and standard LoRA lies not in which parameters are trained, but in how these low-rank parameters are initialised. In standard LoRA, the low-rank matrices are typically initialised with random noise, which introduces uncertainty in the early stages of training and may lead to slower convergence or suboptimal optimisation. By contrast, this research employs singular value decomposition to initialise the original weight matrix, thereby providing a more structured starting point for fine-tuning. This reduces optimisation uncertainty in the early training phase and improves convergence efficiency.

3.3.2 Squared rank (SR) optimising module

Traditional LoRA optimises model training by adding low-rank adapters, but the scaling factor of the adapters in LoRA is tied to the LoRA rank, which can slow down the learning speed and limit performance. This issue is especially pronounced when using higher-rank adapters, potentially causing gradient vanishing problems. To address this, the authors introduce the concept of squared rank in LoRA to mitigate gradient vanishing without increasing inference computational costs, as illustrated in the green block of Fig. 2.

For the feature vectors obtained through SVD Module, based on rank r, the following holds:

WrAB=krWpri=krAB.

After processing with the squared rank factor and SVD module processing, the forward feature is obtained as:

YForward.=Xin(Wres+WrAB)=Xin(U[:,r:]S[r:,r:]V[:,r:]T+krAB).

4 Experiment

4.1 Experimental environment configuration

The proposed model fine-tuning and light-weight multi-agent collaboration system experiment is implemented on the server. The experimental machine is configured as follows: CPU: Intel Xeon Platinum 8370C with 32 cores and 64 threads, base clock 2.8 GHz, turbo boost up to 3.5 GHz; GPU: NVIDIA GeForce RTX 4090; Memory: 256 GB DDR4 ECC RAM. Detailed specifications are listed in Table 1.

4.2 Data set construction

To enable the lightweight expert agent model focusing more effectively on cross-regional collaboration in the construction industry of the GBA the authors collected and compiled academic literature, industry reports, news articles, and records related to the region’s construction sector. Additionally, on-site field investigations were conducted. Subsequently, these raw textual data sources were processed to create a structured fine-tuning data set. The data set construction process primarily involved data preprocessing and the creation of structured data.

4.2.1 Data cleaning and processing

The collected source documents are primarily stored in PDF format and often contain complex formatting structures such as multi-column text, tables, headers, and footers. Manually copying and pasting text directly from these PDFs can lead to content disorder, increased noise, and potentially degrade model performance. Therefore, data preprocessing is required before knowledge training. MinerU is a powerful text mining tool that supports stripping and extracting text from PDFs and saving it in a structured Markdown format (Wang et al., 2024). The authors employ MinerU for data preprocessing. An example of the preprocessing procedure applied to a sample document is shown in Fig. 3.

The text files processed and extracted by MinerU are stored in multiple modality formats. Since this research primarily focuses on the influencing factors of cross-regional collaboration in the construction industry of the GBA, subsequent analysis will be conducted solely on the textual modality data.

4.2.2 Structured-corpus construction

To effectively fine-tune the model, the authors constructed a customised data set targeting the cross-regional collaboration in the construction industry of the GBA, based on the PESTEL analysis framework. After PDFs transformation, the purpose of this step is to convert the text modality knowledge related to cross-regional cooperation stored in Markdown into a form that can be learned by LLMs. It should be noted that due to the long length and dense knowledge of the files, it is time-consuming and laborious to simply rely on manual questioning to build question sets and manually set up answers. Given that LLMs have been proven to possess certain human-like cognitive abilities, the authors use LLMs to generate questions for these texts, supplemented by preset answers and final manual corrections, to achieve efficient data set construction.

Given that some documents are lengthy and there is a limit to the max token for a single input when calling the LLM, the MinerU-processed texts need to be further segmented into smaller chunks. It is important to note that to avoid overly concentrated information in a single segment and the loss of context, the authors set a certain overlap while handling the text segmentation. At the same time, to make the fullest use of the text modality knowledge of the reference literature to construct an effective data set, the batch processing token for the segmented text also needs to be limited. Subsequently, to specifically target the LLM, the authors further developed user input prompts related to the six PESTEL dimensions: Political, Economic, Social, Technological, Environmental, and Legal influencing factors. Ultimately, after LLM and manual correction, the final PESTEL theory framework-based influencing factor consultation corpus was obtained. A sample data set built from the GBA ESG Action Report policy document is summarized in Table 2.

ShareGPT is an instruction fine-tuning data set construction framework based on LLMs. By designing different roles and their corresponding input-output text formats, it enables models to better understand and generate content that aligns with human instruction requirements. To better tailor the trained lightweight expert agent to the specific task scenarios, the authors go beyond just question-and-answer pairs by additionally assigning role definitions to the agent model through system prompts. Finally, following the ShareGPT format, the training data are formatted as: “‘role’: ‘system’, ‘content’: ‘<system prompt>’, ‘role’: ‘user’, ‘content’: ‘<user question>’, ‘role’: ‘assistant’, ‘content’: ‘<reference answer>’,” based on the compiled knowledge Q&A pairs mentioned above. Figure 4 illustrates an example of such annotation.

4.3 Evaluation index

To assess the performance of the proposed knowledge text clustering, evaluation metrics are utilized, which are defined as follows.

4.3.1 Loss

Sigmoid Loss is a commonly used loss function that measures the difference between the model’s predicted probability distribution and the true labels. By monitoring the trend of the Sigmoid Loss, one can determine whether the model is gradually converging toward an optimal state during training. The relevant calculation formula is as follows:

L=−1N∑i=1N[yi⋅ln(y^i)+(1−yi)⋅ln⁡(1−y^i)],

where N represents the number of samples, yi denotes the true label, and y^i denotes the model’s predicted value.

4.3.2 Time cost

Time cost refers to the amount of time required for a model to complete one inference or a single training iteration until convergence. Unlike traditional machine learning and deep learning models, LLMs have a massive number of parameters, resulting in extremely high computational costs per training run. Therefore, achieving convergence as early as possible—while maintaining acceptable accuracy, which means lower computational resource consumption. As the model’s parameter size increases, its cost can grow by more than a hundredfold. Hence, it is essential to monitor the model’s convergence time to select the most suitable training strategy.

4.4 Experiment and results

4.4.1 Base model selection

The selection of an open-source LLM as a base fine-tuning model is necessary. Since LLMs of 7B and larger require high GPU computational power, and under the constraints of equipment conditions, most local deployment training scenarios cannot exceed 7B, while lightweight models of 1.5B, 0.5B, etc., have limited participation in public evaluation experiments, the authors referenced and added different series of 7B-10B parent models as general capability evaluation benchmarks along with the 1.5B foundation. Table 3 shows a comparison of the general capabilities.

Table 3 shows that in the 7-9B scale comparison, although Qwen2.5-7B has lower parameters than Gemma 2-9B, it outperforms or approaches Gemma 2-9B in key indicators such as MMLU (74.2 vs 71.3), MMLU Pro (45 vs 44.7), and GPQA Diamond (36.4 vs 32.8), demonstrating stronger interdisciplinary knowledge understanding and academic reasoning capabilities. At the lightweight model level, Qwen2.5-1.5B achieved a score of 57.6 in Commonsense reasoning tasks, significantly higher than the same-scale Chinese enterprise model DeepSeek-R1-1.5B (44.4), indicating its relative advantage among lightweight models.

Additionally, since this research focuses on cross-regional collaboration in the construction industry under the governance of governments in the GBA, it requires LLMs to possess certain industry-specific basic capabilities, such as administrative and governmental cognition, legal cognition, and certain professional qualification levels. Therefore, referencing existing research released in github (github.com/jeinlee1991/chinese-llm-benchmark/blob/main/leaderboard/MMCU-%E6%B3%95%E5%BE%8B.md), the authors further compared basic capabilities matching the application scenarios of this research. The comparison results are shown in Table 4.

From the comparison results in Table 4, it can be observed that ChatGPT-4o-latest, as a commercial closed-source model, performs optimally in most indicators. Among lightweight models with the same 1.5B parameters, qwen2.5-1.5B-instruct outperforms DeepSeek-R1-1.5B in task evaluations related to administrative affairs, Chinese law, and professional title assessment, achieving 77.8% of the administrative affairs capability performance of the lightweight model gemma-3-4b-it with only 37.5% of its parameter scale. Therefore, among lightweight models, qwen2.5-1.5B-instruct demonstrates certain potential for administrative affairs applications.

In the comparison of medium-to-high-scale parameters from 7B to 30B, for administrative affairs and general legal capability evaluation, qwen2.5-7b-instruct achieved over 80% of the capability of gemma 27B series models with only 25% of the parameter scale in testing, and outperformed gemma series models in civil service selection examinations, verifying its capability for Chinese government-related applications. In senior professional title evaluation, it leads all models including ChatGPT-4o-latest, verifying the professional capabilities of the Qwen series models.

Notably, regarding the MMCU-Law test results, this test set was launched by the Jiaguwen AI Research Institute as a large-scale Chinese multi-task test set, aiming to fill the gap in Chinese large model evaluation benchmarks. It references the widely used MMLU data set in the English domain but focuses on knowledge understanding and reasoning capability evaluation in Chinese contexts. In this test, Qwen2.5-7B scored 48 points, slightly surpassing the ChatGPT-4o-latest model. As a lightweight model, Qwen2.5-1.5B scored 33, significantly outperforming gemma-3-4b-it in the corresponding test (only 21.5), and even exceeding the capability of the non-lightweight model GLM-Z1-32B from the same Chinese origin Z.AI (22.5). This further verifies the adaptive advantages of Qwen series models in Chinese domain tasks within the tri-legal system of the GBA in this research.

Therefore, regarding the comprehensive performance of Qwen series and considering construction industry-related tasks have a strong dependence on Chinese language comprehension, policy text analysis capabilities, and localized deployment, while also requiring the model to be lightweight and easily fine-tuned, finally, the authors selected the Qwen2.5 series’ lightest 1.5B-parameter model as the base model.

4.4.2 Experiment results

After data processing, approximately 8,000 text-label sequences were obtained. In the experiments, the corpus was split into training and validation sets at an 80:20 ratio. Based on the Qwen2.5:1.5B model and the partitioned data, multiple parameter optimisation experiments were conducted. Each test aimed to keep other parameters constant as much as possible, selecting optimal parameters through repeated trials to provide reference for similar future studies. The corresponding training process is illustrated in Fig. 5 and Fig. 6.

Compared to traditional deep learning model training, the secondary fine-tuning of LLMs significantly increases computational resource demands. As the number of model parameters grows, both the cost and time required for training rise substantially. Therefore, selecting appropriate values for the number of epochs and the learning rate directly impacts training efficiency. Regarding the epoch metric for this training, as shown in Fig. 5, the model achieves convergence at 75 and 100 epochs, with relatively good convergence observed at 40 and 50 epochs as well. Since a larger number of epochs corresponds to higher training time and greater computational consumption, and the model’s performance begins to plateau around 40 epochs, choosing 50 epochs strikes a reasonable balance. This choice ensures the model converges while providing an optimal checkpoint for downstream use.

Regarding the learning rate metric, the training of the SVDSR-LoRA model begins to converge around 40 epochs. Therefore, to better capture the model’s performance, and following existing studies, the authors select the training metrics at a loss value of 0.3 as the baseline for evaluating the model’s time cost during training (Smith, 2018). In other words, the evaluation focuses on the time taken for the model to first reach this desired condition. The relevant results are shown in Fig. 6 and Table 5.

Although a learning rate of 3.00E-04 theoretically enables faster training speed, in practice, as training progresses, the time consumption for the model to reach the desired loss is lowest and nearly the same when using learning rates of 7.00E-04 and 8.50E-04.

To better evaluate the stability and convergence characteristics of model training under different learning rate configurations, the authors conducted a comparative analysis of two key indicators after outlier filtering: the gradient norm density distribution and the evolution of gradient norms over training steps. The gradient norm is an important metric that measures the magnitude of gradients during model training, reflecting the intensity of parameter updates. Typically, the gradient norm changes as training progresses. In the initial phase, due to randomly initialised model parameters, the gradient norm can be relatively large, especially with higher learning rates, as the model parameters have not yet converged and gradient directions fluctuate significantly.

During the middle phase, as the model gradually learns patterns from the data, the gradient norm generally stabilizes and decreases to a reasonable range, indicating normal convergence. In the late phase, as the model approaches convergence, the gradient norm tends to further decrease and approach zero because parameter updates become minimal and the model has basically reached an optimal state. By monitoring gradient norms, prolonged excessive gradient vanishing issues can be detected, allowing an assessment of whether model training is progressing appropriately. The statistical properties and dynamic changes of the gradients during training are illustrated in Fig. 7 below.

The density distribution metric in Fig. 7 refers to the probability density distribution of the gradient norm over different numerical ranges. The more concentrated the density peaks and the more they lie within a reasonable range, the more stable the training process. In the experiment, the gradient norm density distribution for the 8.5e-4 learning rate (orange curve) exhibits a distinct peak in the 0.5-0.6 range. This demonstrates better training stability and convergence compared to the broad distribution for the 3.0e-4 learning rate (the red curve extends rightward to the 1.2-1.4 range) and the multi-peak dispersion for the 9.0e-4 learning rate (gray curve). The scatter plot metric in the figure shows the real-time evolution of the gradient norm over the number of training steps. The denser the data points and the more stable the gradient norm values, the more controllable the training process and the less likely it is to encounter exploding or vanishing gradients. In the experiment, with a learning rate of 8.5e-4 (orange dots), the gradient norm remained within a reasonable range of 0.4-0.8 throughout the entire training process of more than 1000 steps. This avoided the occurrence of multiple minor gradient explosions, such as abnormally high values exceeding 1.8 in the gray dots, and the gradient norm continued to rise in the later stages, followed by a lagging decline. This demonstrated good training controllability and long-term stability. Therefore, the 8.5e-4 learning rate is generally considered to be the optimal parameter under 40 epoch.

5 Discussion

5.1 Model evaluation

To assess the effectiveness of the fine-tuned lightweight expert agent model, it is necessary to conduct quantitative and qualitative evaluation of the model’s fundamental general knowledge capability and basic engineering domain capability performance. To provide a more structured and targeted validation, a three-level integrated evaluation framework is adopted. At the foundational level, ROUGE and BLEU serve a diagnostic function in the evaluation that focuses on general Chinese text processing capability, examining whether the lightweight model, after domain-specific fine-tuning, retains its fundamental abilities in Chinese-language understanding, information extraction, and structured generation. On this basis, at the domain level, evaluation focuses on whether the model correctly captures professional knowledge in the construction context, primarily through semantic and content-sensitive metrics such as BERTScore and Token-F1, with ROUGE-L retained only as a supplementary indicator of structural and terminological alignment. Thirdly, at the task level, the effectiveness of the model in identifying influencing factors and generating actionable recommendations is assessed through the coverage of high-frequency factors and expert-based evaluation in terms of relevance, completeness, insightfulness, and practical value.

For the basic cognition validation, the authors utilize the open-source ConstructionQA (huggingface.co/ datasets/ahhany/constructionQAs) and cnewsum-processed (huggingface.co/ datasets/ ethanhao2077/cnewsum-processed) data sets from Hugging Face and based on response semantic similarity metrics as well as ROUGE and BLEU metrics to scored and compared the basic capabilities of multiple fine-tuned lightweight expert models against official models including DeepSeek-Chat, GPT-4o, and Qwen-Plus, to verify the feasibility and effectiveness of the selected model for applications in the engineering domain.

For general-purpose tasks using ROUGE and BLEU metrics, a Chinese news was selected data set to quantitatively evaluate the model’s general Chinese language capability. The results are presented in Table 6.

Overall, except for GPT-4o, a model developed in the United States where English is the native language, which leads significantly on the machine translation metric BLEU, models originating from China show little difference in BLEU performance between full-scale and lightweight versions, and all outperform the Gemma3-1b series model released by Google. Regarding the ROUGE metric, which measures the model’s text comprehension capability, the ROUGE metric tests across all models indicate that full-scale official models consistently outperform lightweight LLMs with parameters below 2B. GPT-4o demonstrates the strongest information coverage capability in the Chinese News summarization task. The fine-tuned lightweight Qwen2.5:1.5b model significantly outperforms the DeepSeek-r1:1.5b model and Gemma3-1b model with similar parameters at both word-level and sequence-level granularities, and can relatively well cover key information in human-written summaries, achieving approximately half the performance level of full-scale commercial models. Notably, the ROUGE-Lsum scores of all models show minimal difference from their ROUGE-L scores, indicating that all models generate summaries with compact structures and good coherence between sentences. Combining BLEU and ROUGE metrics, the fine-tuned Qwen2.5:1.5b model achieves over half the performance of full-scale commercial models despite significant differences in model scale and outperforms both the DeepSeek-r1:1.5b and Gemma3-1b models, which further demonstrates the feasibility of base model selection and the generalisability of the fine-tuned lightweight model.

For measuring fundamental engineering capability, the ahhany/constructionQAs data set open-sourced on Hugging Face was selected as the basic capability evaluation benchmark. The quantitative capability assessment results for different models are presented in Table 7.

In Table 7, BERTScore F1 is used to measure the semantic similarity between the generated answer and the reference text, reflecting the model’s depth of understanding of professional knowledge; ROUGE-L focuses on the overlap with reference answers in ConstructionQA in terms of word order and structure, demonstrating whether the answer organization aligns with standard expression; Token-level F1 measures key information coverage from the perspectives of precision and recall. From the BERTScore F1 perspective, the PESTEL domain fine-tuned Qwen2.5-1.5b outperforms both the DeepSeek-r1-1.5b and Gemma3-1b models in terms of understanding and articulation of professional knowledge on the professional knowledge data set. It also outperforms the official DeepSeek-Chat model in all aspects except for Token-level F1, where it is slightly inferior to the full-scale DeepSeek-Chat model. Notably, since this test is conducted on a data set related to the construction domain, for the same ROUGE testing, the fine-tuned Qwen2.5-1.5b model has achieved 88.26% of the commercial Qwen-Plus full-scale model’s performance in the construction domain, whereas it only reached 54.31% in the general domain, directly and effectively demonstrating the effectiveness of domain-specific fine-tuning for the construction domain. Therefore, after fine-tuning, the Qwen2.5-1.5b model outperforms other lightweight models in both general domain and construction domain and can achieve a relatively comprehensive capability level comparable to conventional full-scale commercial models.

5.2 Training optimising analysis

The above analysis validates the effectiveness of the fine-tuned lightweight expert agent model. However, as mentioned in the previous sections, the inherent constraints of traditional LoRA rank to some extent slow down the learning speed and limit model performance. To address this, the authors introduce a coupled SVDSR module into LoRA training to initialise the original weight matrices and mitigate gradient vanishing for optimising the uncertainty situation in early training and accelerating model convergence. To further evaluate the effectiveness of the proposed SVDSR-LoRA model through the coupled SVDSR module integration, the authors selected the time cost when Loss reaches 0.6 during the middle training period as the speed comparison metric and conducted a direct performance comparison with the standard LoRA method under identical conditions. To better compare the model’s performance before and after introducing the coupled module, comparative experiments were also conducted on both the DeepSeek-r1:1.5B model with the same parameter scale and the Qwen2.5:1.5B lightweight model, under both SVDSR-LoRA and standard LoRA training modes. The results are presented in Fig. 8.

Figure 8(a) demonstrates that on Qwen2.5-1.5B, training with the added coupled SVDSR module, the proposed coupled model, significantly reduces the time to reach the same target loss threshold from 2.80 h to 1.79 h, representing a 36% reduction in time cost. As mentioned earlier, during the intermediate stages of training, as the model gradually learns patterns from the data, the gradient norm generally tends to stabilize and decrease to a reasonable range, indicating gradual convergence. In the later stages, as the model approaches convergence, the gradient norm continues to decrease toward zero, as parameter updates become minimal and the model essentially reaches an optimal state. By monitoring gradient norms, prolonged excessive gradient vanishing issues can be detected, allowing an assessment of whether model training is progressing appropriately. The denser the data points and the more stable the gradient norm values, the more controllable the training process becomes, and the lower the likelihood of gradient explosion or vanishing. From Fig. 8(b), the gradient scatter plot throughout the entire training process with the proposed SVDSR-LoRA method is significantly superior to the standard LoRA approach. The same phenomenon is also observed in the training experiments on DeepSeek-r1:1.5b, which demonstrates the generalisability of this method. However, it is important to note that due to architectural differences between different LLMs, even after training optimisation, there remain differences in time cost performance during training. Since DeepSeek-r1 is a reasoning model with a more complex architecture, training of Qwen2.5:1.5b is significantly faster than that of DeepSeek-r1:1.5b. Consequently, the introduction of the SVDSR coupled module would yield more pronounced optimisation of early and middle training time cost for complex models like DeepSeek-r1:1.5b, reducing it from 5.93 h to 2.92 h, representing a time cost reduction of nearly 50%, which demonstrates the effectiveness of the method.

5.3 Identifying agent-based influencing factors

As mentioned earlier, advancing construction projects requires close collaboration among multiple stakeholders, and a simple single LLM struggles to capture the complexity of issues faced by these diverse parties. When the model size is smaller, limitations in its capabilities lead to more noticeable differences in the quality of its output. To enhance the effectiveness of the lightweight expert agent in identifying influencing factors of cross-regional collaboration, the authors further build a lightweight MAS, according to the trained expert agent specifically for the task of recognizing factors affecting cross-regional collaboration.

The lightweight MAS integrates the trained expert models into a multi-agent group chat framework to simulate real-world scenarios where different experts collaboratively analyze and discuss the influencing factors of cross-regional collaboration in the construction industry of the GBA. Specifically, based on the PESTEL framework, the authors configure a total of eight agents within the “MeetingChat” MAS: one secretary agent, one agent dedicated to summarizing and drafting reports, and six expert agents corresponding to the six domains of Political, Economic, Social, Technological, Environmental, and Legal analysis.

In the constructed multi-agent group chat system framework, each agent is instantiated from a fine-tuned expert model. To ensure that each agent performs its designated role, behavioral guidelines—such as speaking order—are enforced through additional prompt engineering, and each agent was pre-assigning a role description. When a user poses a question, the Secretary Agent acts as the interface with the user and coordinates the discussion among the expert agents. The six PESTEL expert agents then engage in a focused discussion on the relevant factors. Finally, the Report Agent compiles the findings into a report, which is subsequently reviewed by the Secretary Agent together with the user to confirm whether it meets the final requirements. This lightweight MAS is built using Python and the VS Code editor. Specifically, the corresponding pseudocode is provided below.

Finally, in response to the prompt, “Please provide a comprehensive analysis of the influencing factors in cross-regional collaboration within the construction industry of the GBA; and, based on all factors, offer overall recommendations,” the analysis output of the lightweight MAS is shown in Fig. 9 below:

The lightweight MAS is capable of simulating an expert meeting scenario. After thorough discussions among multiple expert agents, the report agent documented and summarized the entire simulated meeting, providing a comprehensive and well-rounded response. The report agent first completed a structured extraction of each analyst’s content and successfully recorded the full set of influencing factors in a markdown-formatted table. Subsequently, the comprehensive report analyst agent analyzed and summarized all the identified factors, highlighting the most significant influences and examining potential interactions among them. Finally, the system offered recommendations and potential implementation plans for promoting cross-regional collaboration in the short, medium, and long-term.

To further substantiate the effectiveness of this system, an additional evaluation of the proposed MAS responses from both qualitative and quantitative perspectives was conducted. Regarding quantitative evaluation, the top 10 most frequently occurring influencing factors across multiple reference documents were compiled and compared with the hit rate of responses generated by the model. The results are presented in Table 8.

Table 8 shows that different models exhibit significant differences in identifying key influencing factors for regional integration. From the hit rate perspective, the proposed lightweight MAS constructed by fine-tuning Qwen2.5:1.5b as the baseline model in this research demonstrates the best performance, achieving a 100% hit rate, significantly outperforming the Qwen2.5:1.5b model itself, which achieved only a 20% hit rate. In comparison, this represents a 4.5-times improvement in accuracy over the base Qwen2.5:1.5b model. Compared to other single models that Qwen2.5:7b with a 70% hit rate and DeepSeek-r1:1.5b with a 60% hit rate, the proposed MAS also maintains a leading position, demonstrating the effectiveness for comprehensive analysis of complex social phenomena.

Notably, Table 8 is not intended as a strict test of external generalisation, but rather as a supplementary assessment of in-domain knowledge coverage. The task addressed in this research is not open-domain prediction, but a bounded and context-specific problem, namely the identification of influencing factors in cross-regional collaboration within the construction industry of the Guangdong–Hong Kong–Macao Greater Bay Area. Therefore, the 100% hit rate should be interpreted as indicating that the constructed MAS is able to recover the high-frequency factors consistently identified across existing studies and policy documents. What this result reflects, therefore, is the degree of alignment with, and coverage of, the current knowledge consensus in the domain, rather than an ability to generalize to independent or previously unseen scenarios. Such a setup may lead to optimistic estimates if interpreted as a generalisation benchmark, and future validation using newly published or temporally separated data would provide a more stringent assessment.

Regarding qualitative evaluation, a Likert scale (1~5) questionnaire was designed based on the metrics of Relevance, Comprehensiveness, and Insightfulness. Evaluations were collected from 22 industry experts and researchers from the engineering and LLM sector. Following the aggregation and averaging of expert scores, the evaluation performance of each model was obtained, as presented in Table 9:

Overall, MAS demonstrates clear superiority across all evaluation dimensions, achieving an average score of 4.03 compared to the 2.5–3.2 range of the other models. Notably, MAS excels in comprehensiveness (4.45) and relevance (4.32), showing its ability to cover multi-dimensional frameworks such as PESTEL while staying closely aligned with the domain context. In contrast, Deepseek-r1:1.5b, though lightweight, outperforms Qwen2.5:1.5b and 7B in several areas because it was trained with reasoning-oriented settings, enabling stronger logical responses. However, the Qwen2.5 series benefits from faster inference speed, making it efficient for certain tasks. MAS, built as a multi-agent system fine-tuned from Qwen2.5:1.5b, not only addresses the reasoning and comprehensiveness gaps but also delivers higher insightfulness (3.50) and practical value (3.73), establishing a significant advantage over all baselines and confirming its effectiveness for research and decision-making applications.

6 Conclusions

Cross-regional collaboration in the construction industry of the GBA faces challenges from external environmental factors such as institutional differences, inconsistent standards, and information asymmetry. Traditional research methods have primarily relied on qualitative human analysis, which is subjective and labor-intensive. Although existing LLMs can meet intelligent application needs, they still lack sufficient vertical domain knowledge due to the low structural complexity of construction-related texts. Considering the intelligent capabilities of LLMs, this research proposes an SVDSR-LoRA fine-tuning approach to explore the construction of lightweight expert agents by fine-tuning existing LLMs with domain knowledge. Using influence factor analysis as a case study, the feasibility of applying lightweight LLMs to intelligent applications in cross-regional collaboration within the GBA’s construction industry was investigated. Experimental results show that the model fine-tuned with the SVDSR-LoRA method begins to converge around 40 epochs and achieves optimal performance at a learning rate of 8.5e-4. Furthermore, based on the fine-tuned lightweight expert agents and the PESTEL framework, a lightweight multi-agent collaborative influence analysis system was developed. This system can simulate a multi-domain expert collaboration environment to reveal external environmental factors affecting cross-regional collaboration in the GBA’s construction industry and provide optimisation recommendations. The findings offer valuable insights for expanding the application of AI technologies in the field of cross-regional collaboration research.

However, this research has some limitations. Despite fine-tuning, the model may still generate “hallucinations” or plausible but incorrect content. There is also a potential risk that the model may inherit or amplify biases present in the training corpus during the training process. Additionally, since this research focuses on the GBA, the generalisability to other regions and non-construction industrial sectors needs to be strengthened. Future work could address these issues by expanding the corpus to include cross-border collaborations worldwide, applying more rigorous prompt constraints, and equipping the model with network verification or third-party data validation tools to mitigate hallucinations, enhance generalisability, and reduce response bias.

References

[1]

Alaghbari W, Al-Sakkaf A A, Sultan B, (2019). Factors affecting construction labour productivity in Yemen. International Journal of Construction Management, 19( 1): 79–91

[2]

Brown T, Mann B, Ryder N, et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33: 1877–1901

[3]

Ding L (2020). Think tank expert recommendations | Ding Lieyun: Supporting New infrastructure and enhancing Old infrastructure. Retrieved Nov. 27, 2025, from biipb.org.cn

[4]

Fang L, (2014). Administrative institutional supply for inter-local cross-regional cooperative governance. Theoretical Investigation, 2014( 1): 19–23

[5]

Göpfert JWeinand J MKuckertz PStolten D (2024). Opportunities for large language models and discourse in engineering design. Energy and AI, 7: 100383

[6]

Hao B W, Liu Y F, Li L Y, Wang J, Peng Y, (2024). The instruction tuning of large language models with multi-modal recommendation instruction. Journal of Beijing University of Posts and Telecommunications, 47( 4): 36–43

[7]

Hassan F U, Le T, Lv X, (2021). Addressing legal and contractual matters in construction using natural language processing: A critical review. Journal of Construction Engineering and Management, 147( 9): 03121004

[8]

He C N, Yu B, Liu M, Guo L, Tian L, Huang J F, (2024). Utilizing large language models to illustrate constraints for construction planning. Buildings, 14( 8):

[9]

He J, Chen Y R, Dai T Y, (2025a). Large language model hallucination reduction strategy based on the rumor propagation mechanism. Shiyan Jishu Yu Guanli, 2025( 2): 96–103

[10]

He JShen YXie R F (2025b). Research on categorical recognition and optimization of hallucination phenomenon in large language models. Journal of Frontiers of Computer Science and Technology, 19(5): 1295–1301

[11]

Houlsby NGiurgiu AJastrzebski SMorrone B, et al. (2019). Parameter-efficient transfer learning for NLP. arXiv:1902.00751

[12]

Hu E JShen Y LWallis PAllen-Zhu ZLi Y ZWang SWang LChen W Z (2021). Lora: Low-rank adaptation of large language models. arXiv:2106.09685

[13]

Hu J H, (2022). The logic of cross-regional public crisis governance and the construction of cooperation mechanism. Social Science Journal, 2022( 2): 50–56

[14]

Kasneci E, Sessler K, Küchemann S, Bannert M, Dementieva D, Fischer F, Gasser U, Groh G, Günnemann S, Hüllermeier E, Krusche S, Kutyniok G, Michaeli T, Nerdel C, Pfeffer J, Poquet O, Sailer M, Schmidt A, Seidel T, Stadler M, Weller J, Kuhn J, Kasneci G, (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103: 102274

[15]

Kumar A, Raghunathan A, Jones R, Ma T Y, Liang P, (2022). Fine-tuning can distort pretrained features and underperform out-of-distribution.

[16]

Lee W J, Lee S K, (2024). Development of a knowledge base for construction risk assessments using BERT and graph models. Buildings, 14( 11): Article 3359

[17]

Li H T, Ai Q Y, Chen J, Dong Q, Wu Z J, Liu Y Q, (2025). Blade: Enhancing black-box large language models with small domain-specific models. In Proceedings of the AAAI Conference on Artificial Intelligence, 39( 2): 24422–24430

[18]

Li X, (2018). Comparative analysis of transregional cooperation mechanisms in Central Eurasia: The silk road economic belt, the Eurasian Economic Union and the “New Silk Road”. Journal of Humanities, 2018( 9): 18–25

[19]

Li X LLiang P (2021). Prefix-tuning: Optimizing continuous prompts for generation. arXiv:2101.0019

[20]

Lv Z H, Chen D L, Lou R R, Alazab A, (2021). Artificial intelligence for securing industrial-based cyber–physical systems. Future Generation Computer Systems, 117: 291–298

[21]

Ma Y, Wen C X, Zhu C S, (2023). Application of cross-regional legislation in territorial planning. City Planning Review, 47( 7): 12–18

[22]

Mahowald K, Ivanova A A, Blank I A, Kanwisher N, Tenenbaum J B, Fedorenko E, (2024). Dissociating language and thought in large language models. Trends in Cognitive Sciences, 28( 6): 517–540

[23]

Meng F, Wang Z, Zhang M, (2024). Pissa: Principal singular values and singular vectors adaptation of large language models. Advances in Neural Information Processing Systems, 37: 121038–121072

[24]

Nedeljković Đ, Kovačević M, (2017). Building a construction project key-phrase network from unstructured text documents. Journal of Computing in Civil Engineering, 31( 6): 04017058

[25]

Pan C D (2024). Analysis of the construction market in the Guangdong–Hong Kong–Macao Greater Bay Area. Retrieved Aug. 11, 2025, from dzb.jzsbs.com

[26]

Pan W A, Chen L, Zhan W T, (2019). PESTEL analysis of construction productivity enhancement strategies: A case study of three economies. Journal of Management Engineering, 35( 1): 05018013

[27]

Pavlik J V, (2023). Collaborating with ChatGPT: Considering the implications of generative artificial intelligence for journalism and media education. Journalism & Mass Communication Educator, 78( 1): 84–93

[28]

Qin X L, Gu X, Li D C, Xu H W, (2025). Survey and prospect of large language models. Jisuanji Yingyong, 45( 3): 685–696 (in Chinese)

[29]

Raza M S, Khahro S H, Memon S A, Ali T H, Memon N A, (2021). Global trends in research on carbon footprint of buildings during 1971–2021: a bibliometric investigation. Environmental Science and Pollution Research International, 28( 44): 63227–63236

[30]

Smith L N (2018). A disciplined approach to neural network hyper-parameters: Part 1–learning rate, batch size, momentum, and weight decay. arXiv:1803.09820

[31]

Su C, Zeng G, Ye L, Xu Y Q, (2021). Influence of trans-regional cooperation innovation on regional diversification in the Yangtze River Delta. Changjiang Liuyu Ziyuan Yu Huanjing, 30( 3): 534–543 (in Chinese)

[32]

Su M H, (2015). Path selection for cross-regional cooperative governance among local governments. Journal of Chinese Academy of Governance, 2015( 5): 57–61

[33]

Sun Y, Zheng Y, Huang H Y, Zhang H, Quan J C, (2024). Multi-loop nested LLM-based multi-agent command and control processes. Journal of Command and Control, 10( 6): 732–739

[34]

Tang C H, Qiu P, Dou J M, (2022). The impact of borders and distance on knowledge spillovers—Evidence from cross-regional scientific and technological collaboration. Technology in Society, 70: 102014

[35]

Varshney D, Atkins S, Das A, Diwan V, (2016). Understanding collaboration in a multi-national research capacity-building partnership: a qualitative study. Health Research Policy and Systems, 14( 1): 64

[36]

Wang BXu CZhao X MOuyang L K, et al. (2024). Mineru: An open-source solution for precise document content extraction. arXiv:2409.1883

[37]

Wright D S (1997). Understanding Intergovernmental Relations (3rd ed.). Baltimore: Paul H. Brookes Publishing

[38]

Wu F, Wang M Q, Li N Y, Zeng Y J, (2021). Research on the factors affecting the cooperation between Hong Kong and Mainland Construction Industry. Journal of Engineering Management, 35( 1): 44–48

[39]

Wu Z L (2023). Hong Kong and Macao join forces to promote the development of the Greater Bay Area. Retrieved August 11, 2025, from cppcc.china.com.cn

[40]

Xiang L Q, Tan Y T, Shen G, Jin X, (2022). Applications of multi-agent systems from the perspective of construction management: A literature review. Engineering, Construction, and Architectural Management, 29( 9): 3288–3310

[41]

Xiang P C, Pang X Y, (2021). Horizontal inter-governmental conflict coordination mechanism for trans-regional major engineering projects. Journal of Beijing Administration Institute, 2021( 3): 42–48

[42]

Xie M J, Qiu Y Z, Liang Y S, Zhou Y K, Liu Z X, Zhang G Q, (2022). Policies, applications, barriers and future trends of building information modeling technology for building sustainability and informatization in China. Energy Reports, 8: 7107–7126

[43]

Xu N, Zhou X Q, Guo C R, Xiao B, Wei F, Hu Y T, (2022). Text mining applications in the construction industry: current status, research gaps, and prospects. Sustainability, 14( 24): 16846

[44]

Xu P Y, Wu G H, (2019). Spatial evolution of the knowledge innovation network in Guangdong–Hong Kong–Macao Greater Bay Area: The role of Shenzhen technological innovation hub. China Soft Science, 2019( 5): 68–79

[45]

Ye QWu T (2024). The total economic output of the Guangdong–Hong Kong–Macao Greater Bay Area exceeds 14 trillion yuan, marking a new level of comprehensive strength. Retrieved Aug. 11, 2025, from gov.cn/lianbo

[46]

Yin Y WCarenini G (2025). ARR: Question answering with large language models via analyzing, retrieving, and reasoning. arXiv:2502.04689

[47]

Yu TDong J (2025). Research on collaborative decision-making of large language models in multi-agent game environments. Computer Engineering

[48]

Yuan W P, Yan X Y, Cao H H, (2018). A study on Yunnan’s regional strategy of connecting “corridors” and “belts” under cross-regional cooperation. Inquiry into Economic Issues, 2018( 2): 130–134

[49]

Yue M R (2025). A survey of large language model agents for question answering. arXiv:2503.19213

[50]

Zhang J S, El-Gohary N M, (2016). Semantic NLP-based information extraction from construction regulatory documents for automated compliance checking. Journal of Computing in Civil Engineering, 30( 2): 04015014

[51]

Zhang M R, (2018). Developments in inter-regional conflict of laws within China. Hong Kong Law Journal, 48: 1097–1135

[52]

Zhang X C, Feng Y W, Xu Y H, (2023). Evolution and interaction between regional cooperation and urban factor agglomeration in the Guangdong–Hong Kong–Macao Greater Bay Area. Zhongguo Renkou Ziyuan Yu Huanjing, 33( 6): 138–150

[53]

Zhang X C, Xia Y H, Shan Z R, Xu S C, (2022). Characteristics and evolution mechanism of inter-governmental cooperation network of Guangdong–Hong Kong–Macao Greater Bay Area. Urban Development Studies, 29( 1): 7–14

[54]

Zou Y, Kiviniemi A, Jones S W, (2017). Retrieving similar cases for construction project risk management using natural language processing techniques. Automation in Construction, 80: 66–76

Rights & permissions

The Author(s)

PDF (12967KB)

487

Accesses

0

Citation

Detail

Sections
Recommended

/

〈 〉