Background: Carbapenem-resistant hypervirulent Klebsiella pneumoniae (CR-hvKp) is a growing concern due to high mortality and limited therapies, with scarce data on novel antimicrobials in vitro activity against it.
Aims: This research evaluated the in vitro efficacy of six novel antimicrobial drugs including ceftolozane/tazobactam, cefiderocol, eravacycline, omadacycline, temocillin, and plazomicin, against CR-hvKP and characterize its molecular epidemiology.
Methods: 106 non-repetitive clinical CR-hvKP strains were collected from Sichuan Provincial People's Hospital between August 2018 and December 2023. CR-hvKP were identified using VITEK-2 Compact and MALDI-TOF MS, confirmed via string tests and PCR. The E-test strip method assessed the in vitro antibacterial activity of novel antimicrobial drugs against CR-hvKP. The molecular characterization of CR-hvKP was conducted by PCR to amplify resistance genes, virulence genes, housekeeping genes, and wzi genes. The Galleria mellonella infection model explored the virulence characteristics of CR-hvKP strains.
Results: CR-hvKP had a relatively high susceptibility rate of 96.2% to cefiderocol, showing good antibacterial activity, whereas ceftolozane/tazobactam, temocillin, omadacycline, eravacycline, and plazomicin exhibited high resistance rates (81.1%–99.1%). ST11-KL64 was the predominant type in CR-hvKP strains. We identified three new ST subtypes, ST8115, ST8116 and ST8117. The most prevalent carbapenemase genes were blaKPC and blaNDM, and approximately 75.5% of CR-hvKP carried blaKPC, blaSHV, and blaCTX-M.
Conclusions: Cefiderocol appears highly promising as a therapy for CR-hvKP infections. Our findings will not only effectively address the challenge of CR-hvKP resistance, but also provide evidence to support the optimization of clinical therapeutic strategies and further promote the development and application of novel antimicrobial drugs.
Objectives: This study aims to provide a comprehensive overview of the cost-effectiveness of COVID-19 vaccination strategies and examine the reporting quality of the included studies.
Methods: A systematic search was performed across PubMed, Web of Science, Embase, Cochrane Library, the International Network of Agencies for Health Technology Assessment, and Chinese databases (CNKI, WanFang, VIP, and SinoMed). Health economic evaluations published from inception to December 19, 2024, considering both costs and effects of vaccination strategies for COVID-19 were included. The reporting quality of the included studies was comprehensively assessed according to the Consolidated Health Economic Evaluation Reporting Standards Checklist (CHEERS) 2022. A narrative synthesis was employed to summarize and present the findings across diverse vaccination strategies. To ensure transparency, the study protocol was registered prospectively in the Prospective Register of Systematic Reviews (CRD42025629706).
Results: A total of 68 studies were included. These studies were heterogeneous concerning reporting quality, vaccination strategies, adopted perspective, applied models, and outcome indicator used. The CHEERS quality scoring rate of 68 studies ranged from 23% to 88%, with a median of 71%. Vaccination was consistently found more cost-effective or cost-saving than no vaccination. Prioritizing high-risk populations, particularly older adults, delivered even higher economic benefits across diverse settings.
Conclusion: COVID-19 vaccination remains cost-effective globally. While rapid deployment and high risk prioritization maximize economic benefits, optimal strategies are sensitive to epidemiological and methodological factors, placing substantial demands on policymakers. Future research should integrate long-term impacts, health equity, and reporting standards to strengthen evidence based vaccination policies.
Background: Umbrella reviews (URs) synthesize evidence across multiple systematic reviews and meta-analyses to inform decision-making. However, evaluating and integrating evidence certainty across overlapping and sometimes conflicting meta-analyses remains a major methodological challenge of URs, limiting the reliability of conclusions.
Objective: To compare evidence evaluation frameworks for URs, assess their suitability for different study types, and provide guidance for managing overlapping evidence, conflicting certainty ratings, and forming overall conclusions.
Methods: We mapped published UR methods to identify current practices and gaps. Frameworks were compared across three dimensions: applicability to study types (interventional, causal observational, and descriptive observational), scalability and reproducibility, and capacity to handle methodological challenges specific to URs. Recommendations were developed through expert consensus and validated via case studies across diverse domains.
Results: Grading of Recommendations Assessment, Development and Evaluation (GRADE) is widely accepted for interventional evidence but shows limited applicability for observational URs due to subjective judgments and reproducibility issues. Credibility frameworks, based on predefined statistical thresholds, offer greater scalability and reproducibility but require adaptation for different study types. Key challenges include: (1) managing overlapping primary studies, (2) resolving conflicting certainty ratings between high-quality reviews, and (3) ensuring transparent evidence integration.
Conclusions: Framework selection should be tailored to study type and context, with credibility frameworks better suited for observational evidence and GRADE optimal for interventional studies. Systematic approaches to overlapping evidence and conflicting ratings are essential for UR validity. We provide practical recommendations for framework selection, strategies for common challenges, and enhanced reporting standards to improve transparency and reproducibility.
Aim: To evaluate efficacy and safety of high-flow nasal cannula (HFNC) versus noninvasive ventilation (NIV) in adults with hypercapnic respiratory failure (HRF).
Methods: We conducted a meta-analysis of randomized controlled trials from PubMed, Embase, The Cochrane Library, Wanfang, CNKI, and Weipu comparing HFNC with NIV in HRF patients. Primary outcomes were mortality and endotracheal intubation. Secondary outcomes included blood gas values, hospital/intensive care unit (ICU) length of stay, nasal skin breakdown, and comfort. Pooled risk ratios (RR) and mean differences (MD) with 95% confidence intervals (CI) were calculated.
Results: Twenty trials involving 1835 patients were included. No significant differences were observed in mortality (RR, 0.94; 95% CI, 0.69 to 1.28), endotracheal intubation (RR, 0.87; 95% CI, 0.71 to 1.06), PaCO2 (MD, –0.87 mmHg; 95% CI, –2.98 to 1.25) or PaO2 (MD, 2.67 mmHg; 95% CI, –0.66 to 6.01). However, HFNC significantly reduced hospital stay (MD, –0.69 days; 95% CI, –1.08 to –0.30) and ICU stay (MD, –0.98 days; 95% CI, –1.50 to –0.45), lowered nasal skin breakdown risk (RR, 0.20; 95% CI, 0.09 to 0.45) and improved comfort (MD, –1.19; 95% CI, –1.98 to –0.41). Subgroup and sensitivity analyses confirmed the principal findings.
Conclusions: HFNC was not significantly different from NIV for mortality and endotracheal intubation in HRF but was associated with shorter hospital and ICU stays, reduced nasal facial skin breakdown and improved comfort. Larger trials in diverse populations are needed.
Aim: This study aimed to evaluate the efficacy and safety of OsteoKing compared to a standard active comparator (Yaobitong capsules) for managing lumbar disc herniation (LDH).
Methods: A randomized, double-blind, double-dummy, positive-controlled, multicenter trial was conducted across 17 hospitals in China. A total of 210 eligible patients with LDH were randomized in a 2:1 ratio to receive either OsteoKing (n = 140) or Yaobitong capsules (n = 70) for 4 weeks. The primary efficacy endpoint was the change in the Visual Analog Scale (VAS) score from baseline to week 4. Secondary endpoints included the Oswestry Disability Index score, the resolution rate of individual traditional Chinese medicine symptoms, and safety assessments. Missing data and repeated measures were handled using multiple imputation and a mixed-model for repeated measures.
Results: Baseline characteristics were well-matched between the two groups. At week 4, both groups experienced substantial improvements in pain and lumbar function. In the full analysis set (FAS), the between-group least squares mean difference (LSMD) in VAS reduction was −0.37 (95% CI: −0.77 to 0.03; p = 0.074). In the per-protocol set (PPS) analysis, the OsteoKing group achieved a significantly greater reduction in VAS scores compared to the control group (LSMD: −0.41, 95% CI: −0.82 to −0.01; p = 0.0453). Improvements in ODI scores were comparable between the groups in both the FAS and PPS analyses (p > 0.05). Additionally, the OsteoKing group demonstrated significantly higher resolution rates for low back pain and fatigue/lassitude at week 4 (p < 0.05). Both treatments were safe and well-tolerated, with comparable incidence rates of adverse events (21.43% vs. 24.29%, p = 0.6396).
Conclusions: OsteoKing demonstrated comparable overall efficacy and safety to Yaobitong capsules in improving lumbar function and relieving symptoms in patients with LDH. Furthermore, OsteoKing may offer a modest, short-term analgesic advantage at the 4-week mark. OsteoKing represents a safe, effective, and viable conservative therapeutic option for LDH.
Objective: To analyze the adherence of Checklist for Artificial Intelligence (AI) in Medical Imaging (CLAIM) in top medical imaging journals.
Methods: A search for AI research in top medical imaging journals was performed from Web of Science Core Collection. The adherence was assessed by the reporting score and compliance rate of CLAIM. Potential influencing factors were also analyzed.
Results: A total of 501 articles were included. After quality assessment, the median CLAIM score was 19 points (25th–75th: 17–21 points), and the median overall compliance rate was 51.4% (25th–75th: 46.2%–57.9%). Among the 42 items, 14 items had compliance rates of ≥80%, 11 items had compliance rates of 40%–79%, while 17 items had compliance rates of <40%. Low compliance rates were predominantly concentrated on items such as quality control of annotations, sample size determination, model robustness evaluation, failure analysis, and code/data availability. In terms of temporal trends, both the overall score (ρ = 0.119, p = 0.008) and compliance rate (ρ = 0.139, p = 0.002) showed weak positive correlations with the publication year. Furthermore, compliance for items 2 (structured abstract) (ρ = 0.129, p = 0.004), 33 (participant flow diagram) (ρ = 0.131, p = 0.003), 34 (demographic/clinical characteristics by partition) (ρ = 0.122, p = 0.006), 41 (study protocol/data/code availability) (ρ = 0.094, p = 0.036), and 42 (funding sources) (ρ = 0.128, p = 0.004) demonstrated significant upward trends over time. Logistic regression analysis identified the following negative predictors of reporting quality: research objectives focused on image reconstruction (OR = 0.416, 95% CI: 0.203–0.842, p = 0.015), artifact reduction (OR = 0.072, 95% CI: 0.009–0.365, p = 0.004), multimodal imaging (OR = 0.301, 95% CI: 0.090–0.899, p = 0.038), use of public (OR = 0.508, 95% CI: 0.289–0.883, p = 0.017) and public–private hybrid (OR = 0.339, 95% CI: 0.146–0.759, p = 0.010) development datasets. Conversely, positive predictors were: publication year (OR = 1.246, 95% CI: 1.067–1.458, p = 0.006); being published in a CLAIM adopting journal (OR = 1.737, 95% CI: 1.070–2.481, p = 0.026); code availability (OR = 1.929, 95% CI: 1.181–3.186, p = 0.009); and development datasets sizes of 100–199 cases (OR = 2.187, 95% CI: 1.121–4.320, p = 0.023), 500–999 cases (OR = 3.409, 95% CI: 1.508–7.902, p = 0.004), and ≥1000 cases (OR = 2.798, 95% CI: 1.335–5.979, p = 0.007).
Conclusions: The adherence of CLAIM among AI research in top medical imaging journals is still inadequate. It is needed to join efforts of researchers, editors, and reviewers to strengthen the application of CLAIM to improve the research quality and reproducibility.
Objective: To evaluate the effectiveness of acupuncture for knee osteoarthritis across different comparators, including usual care, sham acupuncture, waiting-list, pharmacological treatments, and other non-pharmacological interventions.
Methods: An umbrella review of systematic reviews/meta-analyses of acupuncture in adults with knee osteoarthritis was performed. MEDLINE, Embase, Cochrane Library, CNKI, WanFang, and VIP databases were searched. Comparisons were conducted among different acupuncture types. Methodological quality was assessed using Revised Assessment of Multiple Systematic Reviews (R-AMSTAR), and the best available evidence was selected.
Results: Twenty-three systematic reviews were included, with a mean R-AMSTAR score of 30.87. Most acupuncture modalities showed consistent improvements in pain and physical function. Compared with sham acupuncture, usual care, and waiting-list, acupuncture produced clinically meaningful improvements in pain and joint function. Compared with pharmacological treatments, acupuncture demonstrated effects similar to non-steroidal anti-inflammatory drugs, while electroacupuncture was superior in improving joint stiffness and overall response rates. Compared with Tui Na or massage, acupuncture showed a slower onset of effect and similar or slightly inferior improvements in function and stiffness. Adverse events were generally mild and local, with a lower risk of gastrointestinal complications. Treatment effects were most evident at the end of treatment and during short-term follow-up (<3 months), whereas long-term evidence remained limited.
Conclusions: Acupuncture of all modalities relieves pain and improve function in knee osteoarthritis patients, thus is a key non-pharmacological option or add-on therapy when medications are unsuitable. For optimal and sustained effect of acupuncture treatment in knee osteoarthritis, more rigorous sham acupuncture–controlled designs and long-term follow-up are needed.
Background: The routine approach in evidence synthesis of adverse events is to estimate the odds ratio or risk ratio of each individual study and then synthesize the study-specific effects for a pooled average estimate, while seldom consider the potential imbalanced duration of exposures of study arms. This article aims to investigate the potential impact of imbalanced exposure time on harm effects.
Methods: We simulated individual participant time-to-event data based on Cox proportional hazard model, with Weibull function to reshape the distribution of the hazards. We further collapsed the data into aggregated one and fitting both hierarchical Binomial regression model and hierarchical Poisson regression model to estimate the pooled RR and incidence rate ratio (IRR). The percentage bias, mean squared error, and coverage probability were examined.
Results: Our results suggested that imbalanced exposure time between study arms can have substantial impact on the estimation of harm effects in evidence synthesis, especially when the extent of the imbalance exceeds 20%. Estimating an IRR to address the imbalanced exposure time only made sense for non-recurrent events when the between-study heterogeneity is small or moderate. A case study by 22 ongoing trials verified the potential biased estimation when exposure time was imbalanced between study arms.
Conclusions: It is inappropriate to ignoring exposure time when there is a large difference (> 20%) between study arms; while the IRR could be used in some cases, collecting individual participant data for evidence synthesis of adverse events for time-to-event data should be the primary consideration.
Aim: To assess HIV testing coverage over time and identify factors influencing repeated testing and HIV positivity in the general population.
Methods: We employed a mixed-methods design: stratified cluster sampling to obtain HIV testing records and an in-depth interview with staff in primary healthcare institutions to explore drivers of repeated testing variation. Multilevel binary logistic regression models were fitted to explore factors associated with repeated testing and positivity.
Results: A total of 3,010,876 records were included from 2019 to 2022, covering 527,693 participants from 46 primary institutions and 1,375,491 from 26 non-primary departments. The coverage increased from 29.63% (2019) to 36.77% (2022). In primary healthcare institutions, 8.0% of the variance in repeated testing occurred at the institution-level (σ2 = 0.286, p < 0.001). Findings from the interview showed the institutions with higher repeated testing lack an effective deduplication system to identify individuals undergoing frequent testing. In non-primary, departmental-level variation accounted for 10.0% of the total variance in repeat testing (σ2 = 0.367, p < 0.001) and 17.8% in HIV positivity (σ2 = 0.711, p < 0.001). Departments including Respiratory, Hematology, Oncology, Emergency, Dermatology and Venereology, and Infectious Diseases showed higher rates of HIV positivity. Notably, some departments such as Nephrology, Obstetrics and Gynecology were characterized by “high repeated testing but low positivity.”
Conclusion: The strategy achieves a sustained increase in coverage. Future efforts should focus on enhancing information-sharing, reducing unnecessary repeated testing in the general population and strengthening testing in high-yield departments.
The rapid advancement of radiotherapy techniques and systemic anticancer agents has created unprecedented opportunities to improve outcomes for breast cancer patients, while also introducing new challenges related to optimal integration and safety. This consensus, convened by the Breast Cancer Committee of the Chinese Anti-Cancer Association and Chinese Society of Clinical Oncology Breast Cancer Committee, systematically evaluated available literature through November 2025 and employed Delphi methodology to generate evidence-based recommendations for combining radiotherapy with immune checkpoint inhibitors, targeted therapies, antibody–drug conjugates and endocrine agents across disease stages. This work aims to guide clinicians toward safer and more effective integration of radiotherapy with contemporary systemic treatments, promote consistency in clinical practice, and identify priority directions for future research to refine precision radiotherapy and systemic therapy combinations in breast cancer.
Objective: This mixed-method systematic review evaluated the efficacy and safety of Traditional Chinese Medicine (TCM) for metabolic dysfunction-associated steatotic liver disease (MASLD) and identified core TCM herbs/compatibility regimens.
Methods: Six databases were searched (inception to December 31, 2025) for randomized controlled trials (RCTs) of TCM combinations for adult MASLD, with placebo/treatment-as-usual (TAU) as controls. Two independent teams performed screening, data extraction and quality assessment. Quantitative synthesis included meta-analysis with subgroup/meta-regression analyses, publication bias assessment and sensitivity analysis; qualitative synthesis used TCM data mining to identify core herbs and compatibility patterns.
Results: A total of 56 RCTs were included. TCM significantly improved TCM Symptom Score, clinical effective rate, liver controlled attenuation parameter (CAP), liver function (ALT, AST, GGT), lipid metabolism (TC, TG, LDL-C) and reduced BMI, with no benefit for HDL-C. Male proportion had no moderating effect on outcomes. TCM-related adverse events (mainly diarrhea and gastrointestinal discomfort) were mild and rare; 63.2% of studies reported no adverse events. Included RCTs had suboptimal methodological quality (high unclear bias in allocation concealment and blinding). Qualitative analysis of 39 studies identified 78 TCM herbs, with Crataegus pinnatifida (Shan Zha) the most frequent. Six core herbs were identified, with Salvia miltiorrhiza (Dan Shen) and Crataegus pinnatifida (Shan Zha) as the core of the compatibility regimens.
Conclusions: TCM formulations may provide symptomatic and metabolic benefits for MASLD with acceptable safety and identifiable core herbs; however, evidence certainty is very low due to pervasive poor methodological quality of included RCTs, which severely undermines confidence in findings. High-quality RCTs and mechanistic research are warranted.
Aim: Using lung cancer as a model disease, we systematically analyzed the publication characteristics and reporting quality of burden of disease (BoD) studies derived from the Global Burden of Disease (GBD) Database, to provide evidence for standardizing high-quality BoD research.
Methods: A cross-sectional study was conducted. We included GBD-derived lung cancer BoD studies and evaluated reporting quality using 17 core items of the STROBOD Statement. Univariable and multivariable linear regression models were used to explore factors associated with reporting completeness.
Results: A total of 32 studies showed that the number of publications increased significantly since 2025, with repetitive research topics. Evaluation of the 17 core items in the STROBOD Statement revealed that 88.2% (15/17) of the items had reporting deficiencies. The number of reported core items was 11.19 ± 1.67, and none of the included studies reported all of them. The reporting quality of English studies was significantly higher than that of Chinese studies (p < 0.01).
Conclusions: GBD-derived lung cancer BoD studies present severe topic homogeneity and suboptimal reporting quality. Developing tailored reporting guidelines is urgently needed to improve the transparency and reliability of BoD research.
Aim: To develop a guideline for the applicability evaluation tool of the Diagnostic Criteria for Chinese Medicine Syndromes, aiming to enhance the clinical applicability of syndrome diagnostic criteria and promote their widespread adoption and application.
Methods: This guideline was developed using the Delphi method, with reference to the Guideline on Establishing Diagnostic Criteria of Chinese Medicine Syndromes, relevant international evaluation tools, and published diagnostic criteria for Chinese medicine syndromes.
Results: The guideline specifies the principles, procedures, and scoring criteria for evaluating the applicability of traditional Chinese medicine (TCM) syndrome diagnostic criteria. The guideline consists of 5 domains and 25 items: accessibility, readability, implementability, acceptability, and overall evaluation. Each item is assessed on a 5-point Likert scale, with higher scores indicating better applicability.
Conclusions: This guideline provides a systematic and evidence-based framework for evaluating the applicability of diagnostic criteria for Chinese medicine syndromes. It may help improve the clinical applicability of syndrome diagnostic criteria and support their dissemination and implementation in clinical practice.
Objective: Evidence-based medicine emphasizes clinical research driven by important questions, yet Chinese medicine lacks practical quantitative tools to identify such questions. We developed the Chinese Medicine Interventional Clinical Trials Research Question Importance Tool (CMICT-RQIT) to support topic selection and provide transparent criteria for proposal review, promoting high-quality clinical research in Chinese medicine.
Methods: This study followed internationally accepted procedures of conceptualization and operationalization. Using a mixed-methods approach—including a literature review, qualitative interviews, Delphi surveys, expert consensus meetings, and the analytic hierarchy process—we developed the CMICT-RQIT. CMICT-RQIT was prespecified and interpreted as a formative/composite multicriteria decision-support tool.
Results: Following a standardized development process comprising three stages—framework construction, tool development, and tool evaluation—the CMICT-RQIT V1.0 was established. It consists of four domains, 10 facets, and 24 items, each with corresponding composite weights, item explanations, scoring criteria, and an operations manual. Exploratory traditional psychometric analyses showed limited internal consistency and poor structural fit, consistent with the formative/composite nature of the tool and should not be interpreted as evidence that all items measure a single latent trait.
Conclusions: CMICT-RQIT V1.0 can assist researchers in selecting research topics, with its greatest value lying in providing a concise summary of the importance of research question along with a comprehensive 24-item checklist for clinical investigators. At the same time, it offers reviewers transparent criteria for assessing the importance of research questions, thereby promoting the objectivity and fairness of evaluations. CMICT-RQIT V1.0 should be used primarily as a structured checklist and decision-support index.
Meta-analysis with continuous outcomes presents a range of methodological challenges. Among these, two issues have received increasing attention: (i) integrating studies that report only the five-number summary (such as the median, interquartile range, and range) rather than the sample mean and standard deviation (SD), and (ii) accurately quantifying between-study heterogeneity. This review first summarizes recent advances in estimating the sample mean and SD from the five-number summary, covering both normality- and non-normality-based estimation methods. We also review recently developed skewness tests that help determine when normality-based estimators are appropriate and present a practical flow chart for integrating studies with five-number summaries into meta-analysis. Building on this, we discuss methods for quantifying the heterogeneity, focusing on the widely used relative heterogeneity statistic I2 and its limitations, particularly its dependence on study sample sizes. We then review the absolute heterogeneity statistic , which quantifies population-level variation across studies and is invariant to study sample sizes, thus complementing traditional measures. By synthesizing these methodological developments and providing practical guidelines and tools, this review aims to support more rigorous and transparent meta-analytic practice for continuous outcomes, especially in the presence of nonstandard reporting formats and varying degrees of heterogeneity.
Objective: Identifying stroke patients at different disease stages is a prerequisite for clinical research using electronic medical records (EMRs), whereas an artificial intelligence-based model that can be directly applied remains lacking. We therefore develop a large language model (LLM) pipeline for stroke staging model (StrokeSM) in retrospective clinical research.
Methods: StrokeSM was developed using a Chinese national stroke database comprising EMRs from 33,637 patients. A total of 2000 patients were randomly selected from the Tianjin regional stroke database for external validation. StrokeSM comprised three phases: stroke hospitalization identification based on BERT and a bidirectional cross-attention network to fuse present illness history and discharge diagnosis, symptom–time extraction based on chief complaint through a UIE-base LLM, and stroke staging classification according to the predefined rules.
Results: On the test set, StrokeSM achieved accuracy, F1 score, precision, and recall of 0.90, 0.91, 0.91, and 0.90, respectively. The F1 score, precision, and recall of StrokeSM for acute phase was 0.91, 0.89, and 0.93, respectively. On the external validation set, StrokeSM had an accuracy, F1 score, precision, and recall of 0.92, 0.93, 0.93, and 0.92, respectively. Moreover, StrokeSM performed remarkably well in acute phase, with F1 score, precision, and recall of 0.97, 0.98, and 0.96, respectively.
Conclusions: StrokeSM had achieved state-of-the-art performance, providing an accurate method of classifying stroke populations with different disease stages in EMRs, especially in the acute phase. StrokeSM heralds automatic and accurate identification of disease stage phenotypes based on LLM in EMRs, laying the foundation for drawing reliable conclusions in clinical research.