BPG is committed to discovery and dissemination of knowledge
Prospective Study Open Access
Copyright: ©Author(s) 2026. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution-NonCommercial (CC BY-NC 4.0) license. No commercial re-use. See permissions. Published by Baishideng Publishing Group Inc.
World J Psychiatry. Oct 19, 2026; 16(10): 123904
Published online Oct 19, 2026. doi: 10.5498/wjp.123904
Feasibility of human-in-the-loop multimodal generative artificial intelligence chatbot for school-based adolescent mental health support
Xi-Wang Fan, Ting-Yu Lin, Li-Cheng Wu, Xin-Xi Chen, Jun-Yao Chen, Yi-Zhen Wu, Yu Zheng, Hui Zhao, Clinical Research Center for Mental Disorders, Shanghai Pudong New Area Mental Health Center, School of Medicine, Tongji University, Shanghai 200124, China
Ming-Hao Wang, Tang Li, Department of SituTech AI Lab, Beijing Situ Wellbeing Technology Co., LTD, Beijing 100080, China
Hao Du, Yao-Xuan Wang, Yu-Jie Wang, Jing-Wen Zhang, Department of General Manager, Beijing Ruihong Embodied Intelligence Robot Technology Co., Ltd., Beijing 100176, China
Yao-Xuan Wang, Yu-Jie Wang, Jing-Wen Zhang, Department of Artificial Intelligence, Beijing Ruihong Embodied Intelligence Robot Technology Co., Ltd., Beijing 100176, China
Li-Wei Huang, Department of Management, Beijing Gaidi Intelligent Technology Co., Ltd., Beijing 100070, China
Tian-Ze Fan, Department of Radiology, The 72nd Group Army Hospital of the Chinese People's Liberation Army, Zhejiang 313100, China
Bing Liu, Department of Radiology, Chinese PLA Medical School, Beijing 100853, China
Bing Liu, Shi-Jun Li, Department of Radiology, The First Medical Center of Chinese PLA General Hospital, Beijing 100853, China
Jie Liu, Beijing Key Laboratory of Design and Intelligent Machining Technology for High Precision Machine Tools, Beijing University of Technology, Beijing 100124, China
Pei Sun, Faculty of Health and Wellness, City University of Macau, Macao 999078, China
Pei Sun, Tsinghua Laboratory of Brain and Intelligence and Department of Psychological and Cognitive Science, Tsinghua University, Beijing 100084, China
Guo-Lin Ma, Department of Radiology, China-Japan Friend Hospital, Beijing 100029, China
Qiang Cheng, College of Mechanical and Energy Engineering, Beijing University of Technology, Beijing 100124, China
ORCID number: Xi-Wang Fan (0000-0003-4180-0496); Ting-Yu Lin (0009-0003-2171-1565); Li-Cheng Wu (0009-0009-2889-4935); Xin-Xi Chen (0009-0003-4332-999X); Jun-Yao Chen (0009-0009-5337-5002); Yi-Zhen Wu (0009-0001-5007-0336); Yu Zheng (0009-0003-5325-8080); Pei Sun (0000-0002-5938-8180); Hui Zhao (0009-0007-4100-0655); Shi-Jun Li (0000-0001-8185-0549).
Co-first authors: Xi-Wang Fan and Ting-Yu Lin.
Author contributions: Fan XW and Lin TY contribute equally to this study as co-first authors; Fan XW, Lin TY, Wang MH, Zheng Y, Chen JY, and Li SJ conceived and designed the study; Fan XW, Lin TY, Wang MH, Wu LC, Li T, Du H, Wang YX, Wang YJ, Zhang JW, Huang LW, Fan TZ, Liu B, Zhao H, and Li SJ contributed to study implementation, including participant recruitment, on-site coordination, technical support, and data acquisition; Chen XX and Lin TY conducted the statistical analyses; Lin TY, Chen XX, and Chen JY verified the analytic outputs; Fan XW, Lin TY, Wang MH, Chen XX, Wu YZ, Wu LC, Chen JY, Zheng Y, and Li SJ drafted the manuscript; all authors reviewed and revised the manuscript critically for important intellectual content and approved the final version for publication; authors affiliated with technology companies contributed to platform development, technical operations, technical troubleshooting, and/or implementation support according to their roles and they did not have sole authority over outcome definition, inferential statistical analysis, interpretation of clinical outcome findings, or manuscript conclusions; Fan XW and Li SJ accept responsibility for the work as a whole, including the integrity of the data, the accuracy of the analyses, and the transparency of role separation.
AI contribution statement: During revision, the authors used ChatGPT (OpenAI, GPT-5.6 SOL; web application; accessed in July 2026) solely for English-language editing of text drafted by the authors in the manuscript and the point-by-point response. The tool was used to improve grammar, sentence structure, word choice, concision, and readability. It was not used to generate scientific content or substantive responses to the reviewers, search for or synthesize literature, generate or verify references, generate or modify study data, conduct statistical analyses, interpret the results, or draw conclusions. The authors independently developed all scientific content and substantive responses, reviewed and approved every AI-assisted language edit, and take full responsibility for the accuracy, integrity, and originality of the manuscript and the point-by-point response.
Supported by the Beijing Nova Programme Interdisciplinary Cooperation Project, No. 20240484674; Capital’s Funds for Health Improvement and Research, No. CFH2024-2-5024; and the Tongji University Medicine-X Interdisciplinary Research Initiative, No. 2025-0708-YB-02.
Institutional review board statement: This study was approved by the Ethics Committee of Chinese PLA General Hospital (Approval No. S2024-849-01).
Clinical trial registration statement: This study was registered at the Chinese Clinical Trial Registry (https://www.chictr.org.cn/). The registration identification number is ChiCTR2500100806.
Informed consent statement: All participants and their guardians provided informed consent before participation.
Conflict-of-interest statement: The authors declare potential competing interests related to professional affiliations and platform development. Ming-Hao Wang and Tang Li are employees of Beijing Situ Wellbeing Technology Co., Ltd., which contributed to the development of the Duoduo system. Hao Du, Yao-Xuan Wang, Yu-Jie Wang, and Jing-Wen Zhang are employees of Beijing Ruihong Embodied Intelligence Robot Technology Co., Ltd., and Li-Wei Huang is an employee of Beijing Gaidi Intelligent Technology Co., Ltd.; these authors contributed to technical operations, platform implementation, and/or study support activities. To reduce the risk of interpretive or commercial bias, role separation was implemented throughout the study. Company-affiliated authors contributed to system development, technical deployment, troubleshooting, and platform-related implementation support, but they did not have sole authority over participant eligibility criteria, outcome selection, inferential statistical analysis, interpretation of clinical outcome findings, or manuscript conclusions. Clinical outcome analyses were conducted by the named analysis team, and analytic outputs were checked by the named verification team. The guarantor authors take responsibility for the integrity of the data, the accuracy of the analyses, and the decision to submit the manuscript. Access to identifiable participant information, raw interaction logs, multimodal records, and monitoring dashboards was restricted according to role-based permissions required for technical operation, safety monitoring, and school follow-up. Company-affiliated authors did not independently use raw interaction logs or multimodal records to define clinical outcomes, adjudicate treatment effects, or determine the interpretation of the reported findings. The Duoduo system was used in this study as a supervised school-based research and implementation tool. No participant fees were charged for use of the system during the study, and the reported clinical outcome analyses were not linked to sales, marketing, or participant-level commercial decisions.
CONSORT 2010 statement: The authors have read the CONSORT 2010 Statement, and the manuscript was prepared and revised according to the CONSORT 2010 Statement.
Data sharing statement: Raw interaction logs and multimodal records will not be shared publicly because they may contain identifiable or sensitive information from minors. Any secondary use of de-identified analytic data or supporting materials will require appropriate ethics approval, privacy protection, and institutional data-governance review, and will not include access to identifiable raw audiovisual or interaction records.
Corresponding author: Shi-Jun Li, Department of Radiology, The First Medical Center of Chinese PLA General Hospital, No. 28 Fuxing Road, Haidian District, Beijing 100853, China. shijunli07@yeah.net
Received: June 2, 2026
Revised: July 7, 2026
Accepted: September 8, 2026
Published online: October 19, 2026
Processing time: 132 Days and 0.2 Hours

Abstract
BACKGROUND

Schools often identify signs of emotional distress in adolescents before they access formal mental health services; however, many school counseling units lack the capacity to provide individualized support. Generative artificial intelligence (GenAI) chatbots may offer an accessible, low-threshold means of providing support. Nevertheless, their use among minors requires staff supervision, strong privacy protections, and clearly defined protocols for identifying and escalating risk-related content.

AIM

To evaluate the feasibility, user engagement, working alliance, safety monitoring workflow, and exploratory short-term outcome indicators of Duoduo, an acceptance and commitment therapy-informed multimodal GenAI chatbot developed for adolescents with elevated depressive symptoms.

METHODS

We conducted a prospective, non-randomized, controlled pilot study at a junior high school in Zhejiang Province, China. Students with elevated depressive symptoms were assigned to either an intervention group (n = 30) or a waitlist control group (n = 32). The intervention comprised eight supervised sessions delivered over two weeks. Feasibility outcomes included completion, engagement, working alliance, and predefined safety-monitoring indicators. Exploratory clinical outcomes were evaluated using baseline-adjusted analysis of covariance, with false discovery rate (FDR) correction applied for multiple comparisons.

RESULTS

Among the 30 students who initiated the intervention, 26 (86.7%) completed all scheduled sessions. Participants used the system for a mean total of 101.3 minutes. The mean Working Alliance Inventory-Short Revised score was 3.56 (SD = 0.86). Nine students triggered dangerous-behavior alerts, and school counselors conducted offline follow-up according to the established safety protocol. In the unadjusted analyses, nominal between-group differences were observed for anxiety and optimism. However, these differences were no longer evident after adjustment for baseline symptom severity and did not remain significant after FDR correction for multiple comparisons. No clinical outcome demonstrated a statistically significant between-group difference after adjustment.

CONCLUSION

Supervised school-based implementation of the chatbot was feasible and supports further evaluation in randomized controlled trials to determine its clinical effectiveness and safety.

Key Words: Generative artificial intelligence; Adolescent mental health; Chatbot; Feasibility study; Depression; Nonrandomized pilot study; Quasi-experimental study; Acceptance and commitment therapy; Safety monitoring; Human-in-the-loop

Core Tip: This pilot study evaluated the supervised school-based implementation of Duoduo, an acceptance and commitment therapy-informed multimodal generative artificial intelligence chatbot developed for adolescents with elevated depressive symptoms. Most participants completed the 2-week intervention, engaged with the system for an average total duration of approximately 101 minutes, and reported a moderate working alliance. Risk-related conversation signals were referred for staff review through a predefined human-in-the-loop safety workflow. The findings primarily support the feasibility of implementation and provide guidance for future program planning. Clinical outcome findings were considered exploratory because participant allocation was pragmatic, baseline symptom severity differed between groups, multiple outcomes were evaluated, and the follow-up period was brief.


  • Citation: Fan XW, Lin TY, Wang MH, Wu LC, Chen XX, Chen JY, Wu YZ, Zheng Y, Li T, Du H, Wang YX, Wang YJ, Zhang JW, Huang LW, Fan TZ, Liu B, Liu J, Sun P, Zhao H, Ma GL, Cheng Q, Li SJ. Feasibility of human-in-the-loop multimodal generative artificial intelligence chatbot for school-based adolescent mental health support. World J Psychiatry 2026; 16(10): 123904
  • URL: https://www.wjgnet.com/2220-3206/full/v16/i10/123904.htm
  • DOI: https://dx.doi.org/10.5498/wjp.123904

INTRODUCTION

Adolescent mental health is a significant public health concern, and existing support systems remain insufficient to meet the growing demand for care. Globally, many children and adolescents with mental disorders do not receive appropriate treatment, with particularly limited access to care for depression and anxiety[1]. In China, high academic pressure, limited mental health literacy, and restricted availability of psychological services contribute to a substantial treatment gap among adolescents. Recent studies have demonstrated that depressive symptoms are prevalent among Chinese children and adolescents, while access to appropriate mental health care remains inadequate[2,3]. Furthermore, depressive symptoms during adolescence commonly co-occur with anxiety, and rumination has increasingly been recognized as a transdiagnostic process associated with anxiety and depression. These findings highlight the need for accessible interventions that address broader emotional difficulties rather than focusing solely on depression[4].

Schools are an important setting for adolescent mental health support because they provide consistent access to young people within their everyday environments. However, school-based mental health services often face significant capacity limitations, and adolescents’ engagement with these services may be affected by practical constraints, stigma, confidentiality concerns, and individual help-seeking preferences[5]. In many Chinese secondary schools, access to individualized psychological support remains limited, and stigma associated with seeking mental health assistance may further hinder service utilization[6,7]. These challenges highlight the need for scalable, low-threshold interventions that can be safely integrated into routine school-based support systems.

Digital mental health interventions may improve accessibility and reduce stigma; however, earlier-generation tools have often been limited by low engagement, insufficient personalization, and high attrition rates[8,9]. Youth-focused digital interventions have demonstrated feasibility and potential clinical benefits, but their effectiveness in real-world settings depends largely on engagement, delivery context, and implementation design[10]. Recent advances in generative artificial intelligence (GenAI) have introduced new opportunities for more naturalistic, context-sensitive, and empathic interactions, although evidence regarding their application among adolescents remains limited and heterogeneous[11,12]. Such responsive interactions may be particularly valuable for adolescents, who may experience difficulties in explicitly expressing emotional distress and may benefit from immediate, nonjudgmental conversational support.

Acceptance and commitment therapy (ACT) is particularly suitable for brief digital delivery because it emphasizes psychological flexibility, present-moment awareness, cognitive defusion, acceptance, and values-guided action[13,14]. Instead of focusing solely on symptom reduction, ACT aims to help individuals modify their relationship with distressing internal experiences while continuing to pursue personally meaningful goals. This process-based framework is relevant for adolescents experiencing anxiety, rumination, and low optimism, and meta-analytic evidence suggests that ACT is a promising approach for addressing adolescent anxiety and depressive symptoms[15]. Furthermore, the modular structure of ACT facilitates structured implementation through chatbot-based interventions.

Most existing artificial intelligence mental health chatbots primarily rely on text-based interactions[9,16-18]. Therefore, they may overlook important aspects of adolescents’ expression of distress, including facial expressions, vocal tone, hesitation, posture, and other nonverbal cues[19,20]. Multimodal systems may, in principle, enable a more comprehensive understanding of contextual information during supportive dialogue; however, whether these cues can be captured reliably and applied in a clinically meaningful manner among adolescents remains an open question. This pilot study can only begin to explore this potential rather than provide definitive evidence. Such approaches should be implemented only when robust privacy protections, data minimization practices, and human oversight are incorporated into the implementation framework.

We developed Duoduo to investigate this question within a school-based context. Duoduo is a supervised multimodal GenAI chatbot that integrates verbal input and selected nonverbal emotional cues to support a brief ACT-informed conversational program. It was implemented as an adjunct school mental health support tool, rather than as a diagnostic system, crisis intervention, or substitute for professional judgment.

Evidence on GenAI chatbots for adolescents is increasing; however, evidence remains limited regarding supervised, school-based, multimodal systems with predefined risk-escalation pathways. This specific configuration, rather than chatbot-based support in general, represents the knowledge gap addressed in this study. Therefore, we designed this pilot study to examine implementation feasibility and oversight procedures within a real-world school setting. We evaluated whether sessions could be delivered during the school day and whether students would complete and engage with the program. We also assessed whether staff could review risk-related content within the established workflow and whether short-term outcome trends supported the need for future randomized trials. A nonrandomized pragmatic design was adopted because randomization was incompatible with the school timetable, which determined student availability during the fixed intervention period. Given the pragmatic allocation approach and brief follow-up period, all clinical analyses were considered exploratory.

MATERIALS AND METHODS
Study design

This prospective, nonrandomized pilot controlled study with a waitlist control group was conducted in a junior high school in Zhejiang Province, China. The primary aim was to evaluate whether a school-based multimodal GenAI chatbot could be implemented, used, and safely monitored among adolescents with elevated depressive symptoms. This study was not designed to assess treatment efficacy.

This study was approved by the Ethics Committee of Chinese PLA General Hospital (Approval No. S2024-849-01) and registered with the Chinese Clinical Trial Registry on April 15, 2025 (ChiCTR2500100806). After registration, study-specific recruitment, informed consent from guardians and participants, enrollment, group allocation, baseline assessments, and intervention procedures were conducted. Due to practical constraints related to the school schedule, participants were pragmatically assigned rather than randomized. Group allocation was primarily determined by students’ availability to attend sessions during a fixed lunch-break period and their willingness to initiate the intervention immediately.

The intervention group completed the chatbot program over 2 weeks, whereas the waitlist control group received no intervention during this time and was provided access to the program after completion of the post-intervention assessment.

Participants and Recruitment

In May 2025, 1287 first-year junior high school students were screened for depressive symptoms. Students were eligible for inclusion if they scored 20 or higher on the Center for Epidemiologic Studies Depression Scale (CES-D) and expressed willingness to participate. Among the 93 students who met this symptom-risk threshold, 64 provided consent and were pragmatically assigned to either the intervention group (n = 32) or the waitlist control group (n = 32) based on school scheduling constraints and immediate availability. Two students assigned to the intervention group withdrew before initiating the intervention, received no treatment, and did not complete the baseline assessment. Therefore, they were excluded from all analyses. The final analytic sample comprised 62 participants, including 30 students who initiated the intervention and 32 waitlist controls.

Among the 30 students who initiated the intervention, 26 (86.7%) completed all eight scheduled sessions. The remaining four students (13.3%) voluntarily discontinued the full session sequence but completed the post-intervention assessment. The study flowchart is shown in Figure 1. Feasibility and safety outcomes were reported using the appropriate denominator for each indicator.

Figure 1
Figure 1 Participant flow diagram illustrating screening, pragmatic allocation, follow-up, and analysis. This figure shows the participant flow, including the number of students screened, assessed for eligibility, pragmatically assigned to the intervention or waitlist control group, lost to follow-up, and included in the final analyses.
Intervention platform, multimodal processing, and delivery workflow

Duoduo is a multimodal virtual agent designed as a structured adjunct to school-based mental health support for adolescents. The platform integrates multimodal signal processing, speech recognition and synthesis, a staged ACT-informed session framework, and dialogue generation optimized for empathic and developmentally appropriate interactions. The model operation diagram is illustrated in Figure 2.

Figure 2
Figure 2 Human-in-the-loop architecture and supervised safety workflow of the Duoduo system. Schematic illustration of how verbal, visual, and vocal inputs were processed into a summarized emotional context to support acceptance and commitment therapy -guided dialogue generation and response delivery, alongside the human-led risk review and escalation workflow (Supplementary material). RAG: Retrieval-augmented generation; LLM: Large language model; ACT: Acceptance and commitment therapy.

User input consisted of verbal and nonverbal signals. Visual, vocal, and textual data were processed to generate a summarized emotional context that incorporated semantic-emotional content and selected nonverbal indicators, including facial expression, head and eye movements, and valence-arousal estimates[20]. To minimize data exposure, nonverbal signals were processed in near real time and transformed into aggregated emotional features rather than stored or used as continuous raw audiovisual streams for clinical decision-making.

The intervention followed a predefined session structure for all participants. At each session stage, the chatbot generated supportive conversational responses combined with ACT-informed guidance. Responses were delivered through text, emotionally expressive synthesized speech, and an animated avatar. In this pilot study, the chatbot operated exclusively within a staff-supervised school workflow, and risk-related content was directed to a separate alert-and-review pathway. To ensure consistent system behavior, the generative model, model version, staged ACT system prompts, and safety guardrails remained unchanged throughout data collection. These components were not updated, retrained, or otherwise modified during the study period. Dialogue generation used Qwen (qwen-max-2025-01-25; Alibaba), accessed through the Qwen application programming interface.

Safety monitoring, governance, and data handling

Safety procedures were based on staff-led review rather than autonomous chatbot decision-making. The chatbot did not provide diagnoses, assess suicide risk, or deliver crisis management recommendations. After each session, Qwen analyzed the complete conversation transcript using semantic recognition to identify predefined dangerous behavior categories, including suicidal or self-harm ideation, thoughts of death, self-injury, bullying victimization, and family violence. When any predefined category was detected, the system generated an alert for authorized staff. Notifications were sent through an in-app message and a short message service (SMS) notification. The May 2025 deployment used a single binary alert criterion (dangerous behavior detected vs not detected) and did not incorporate graded risk stratification. Staff reviewed flagged content and determined whether school-based follow-up was required according to the approved study protocol. Operationally, each alert was managed through a standardized review-and-response pathway. The notified staff member reviewed the flagged conversation and assessed whether it represented a potential safety concern. When follow-up was considered necessary, staff conducted an in-person discussion with the student within the school setting. Where appropriate, staff involved the student’s guardian and the school counseling service according to standard duty-of-care procedures.

Before data collection, school counselors and authorized study staff completed structured training on the monitoring and follow-up procedures. The training covered the use of the monitoring dashboard, interpretation of automatically flagged content, and the alert-review and escalation pathway. It also included criteria for school-based follow-up, available referral options, and confidentiality and data governance requirements for identifiable information. Only staff who had completed this training were authorized to review alerts and conduct follow-up procedures.

Access to monitoring interfaces and identifiable participant data was restricted to authorized personnel responsible for study oversight and school-based follow-up. Study identifiers were stored separately from questionnaire responses and interaction records. Multimodal inputs were transformed into summarized contextual features to support response generation and alerting functions. plementation process rather than as evidence of validated risk-detection performance[21,22].

This study implemented role separation to minimize potential conflicts of interest. Company-affiliated authors contributed to platform development, technical maintenance, and on-site implementation support, whereas academic and clinical investigators were responsible for study design, recruitment procedures, outcome selection, statistical analysis, interpretation of results, and manuscript conclusions. Access to identifiable interaction data, monitoring dashboards, and safety alerts was managed through role-based permissions and limited to personnel involved in platform operation, safety oversight, or school-based follow-up. Clinical datasets were prepared and analyzed by the designated analysis team, and all analytic outputs were reviewed by the verification team. Raw interaction logs and multimodal records were not used by company-affiliated authors to independently define clinical outcomes, conduct inferential statistical analyses, or influence the clinical interpretation of study findings.

Raw audiovisual data were used solely to generate session-level emotional context summaries and were not considered independent diagnostic evidence. Access to identifiable information, interaction logs, and alert dashboards was controlled through role-based permissions. Data storage, retention, and any secondary use complied with the approved ethics protocol and institutional data governance requirements. Alerts involving potential self-harm or suicide-related content were reviewed by authorized staff according to a predefined escalation procedure, and the chatbot was not authorized to independently assess clinical risk or make crisis management decisions.

Intervention

Participants in the intervention group completed an eight-session program delivered over eight school days within 2 weeks, with sessions conducted during the lunch break. Each session was conducted in a private counseling room using a dedicated tablet. Sessions lasted approximately 15-20 minutes and included a brief mindfulness guidance video (approximately 3-5 minutes) followed by a 10-15-minute ACT-informed chatbot interaction.

The intervention followed a staged ACT framework. Sessions 1-3 focused on awareness and cognitive defusion, sessions 4-6 emphasized present-moment awareness and values clarification, and sessions 7-8 focused on committed action and personal growth.

To support continuity across sessions, recent conversation history was incorporated with each student’s current input during dialogue generation. Interaction logs and automated alerts were reviewed as part of the study oversight process, and risk-related alerts were communicated to designated school staff when follow-up was considered necessary (Figure 3). The system functioned as a structured digital support tool within the school environment and was not designed to replace clinical assessment, crisis management, or professional judgment.

Figure 3
Figure 3 Selected user-facing and monitoring interfaces of the Duoduo platform. A and B: The images show the user-facing chat interface and emotional feedback display; C: The intervention stage selection interface; D: The monitoring interface used to flag risk-related content for staff review (Supplementary material).
Measures

Since this was a pilot study, the primary outcomes were feasibility, engagement, working alliance, and safety monitoring processes. Feasibility was assessed using indicators including completion of all eight scheduled sessions, number of messages exchanged, total duration of use, and activation of risk-related alerts. Because this was an early feasibility pilot, formal success thresholds were not preregistered; therefore, feasibility indicators were reported descriptively. For interpretive purposes, observed values were compared with reference levels commonly used in pilot studies (e.g., session completion and retention rates of approximately 75%) rather than being evaluated as predefined pass/fail criteria. Working alliance was assessed among participants in the intervention group using the Working Alliance Inventory-Short Revised (WAI-SR)[23,24]. Acceptability was not evaluated using a dedicated satisfaction or usability scale; instead, it was inferred from completion rates, engagement metrics, and working alliance indicators.

Clinical outcomes were considered exploratory. The primary clinical domains assessed were depressive symptoms, anxiety symptoms, and rumination. Depressive symptoms were measured using the CES-D[25,26], with the Patient Health Questionnaire-9 (PHQ-9)[27,28] included as a supplementary sensitivity measure. Anxiety symptoms were evaluated using the Generalized Anxiety Disorder-7 (GAD-7)[29,30], and rumination was assessed using the Rumination Response Scale (RRS)[31]. Secondary outcomes included psychological inflexibility measured with the Acceptance and Action Questionnaire-II (AAQ-II)[32], emotion regulation assessed using the Emotion Regulation Inventory (ERI)[33], and optimism evaluated using the Life Orientation Test-Revised (LOT-R)[34,35]. All measures were completed at baseline and post-intervention (2 weeks). All questionnaires were self-report instruments administered to participants in both groups within the school setting. Trained study staff supervised questionnaire administration using standardized instructions. The same instruments were administered at baseline and at the 2-week post-intervention assessment.

Feasibility and engagement metrics

Feasibility, working alliance, engagement, and safety-monitoring indicators were reported using the most appropriate denominator for each measure. Program completion was defined as completion of all eight scheduled sessions and was reported conservatively as the proportion of participants assigned to the intervention group. Message counts and total duration of use were summarized among participants with available system usage data.

Analysis of population and missing data

Two students assigned to the intervention group withdrew before initiating the intervention, did not complete the baseline assessment, and were excluded from all analyses. Therefore, baseline characteristics and outcomes were reported for the analytic sample of 62 participants (30 intervention, 32 control). All 30 intervention participants who initiated the program completed both baseline and post-intervention assessments, as did all 32 control participants. Accordingly, no post-intervention outcome data were missing, and no imputation was required. Linear mixed models were used to account for the repeated-measures study design. Among the 30 intervention participants, four (13.3%) voluntarily discontinued the full session sequence but completed the post-intervention assessment. To assess whether session non-completion was systematic, baseline demographic and clinical characteristics were compared between these four non-completers and the 26 completers (Supplementary Table 1).

Statistical analysis

Since this study was nonrandomized, the groups differed in baseline symptom severity, and the follow-up period was limited to two weeks, clinical analyses were considered descriptive and estimation-oriented rather than confirmatory. Baseline characteristics, feasibility and engagement indicators, working alliance scores, and safety procedure activations were summarized using standard descriptive statistics.

For the clinical measures, the intervention group had greater baseline symptom severity than the control group; therefore, raw post-intervention means were not directly comparable. The primary analysis used analysis of covariance (ANCOVA), in which each post-intervention score was modeled as a function of group assignment and adjusted for the corresponding baseline score, age, and sex. The adjusted between-group difference (intervention minus control) with its 95% confidence interval (95%CI) was reported. Baseline adjustment directly accounts for measured baseline imbalances and may reduce, but cannot eliminate, the influence of regression to the mean. Consistent with the original analysis plan, linear mixed models with a participant-level random intercept and fixed effects for time, group, and their interaction were also fitted, with adjustment for age and sex. These models were used to estimate within-group changes over time. Because 9 outcomes were assessed, the false discovery rate (FDR) across between-group comparisons was controlled using the Benjamini-Hochberg procedure. Given the small, unpowered feasibility sample and nonrandomized design, all P values were interpreted descriptively, and all clinical findings were considered exploratory. Covariate adjustment cannot eliminate selection bias resulting from pragmatic allocation. Propensity-score matching was not applied because it would produce unstable estimates in a sample of this size. Therefore, baseline-adjusted ANCOVA was selected as the more stable approach for addressing baseline imbalance.

RESULTS
Participant flow and baseline characteristics

In May 2025, 1287 first-year junior high school students completed screening for depressive symptoms. Among these students, 93 met the predefined symptom-risk threshold, and 64 provided consent to participate. Participants were pragmatically assigned to either the intervention group (n = 32) or the waitlist control group (n = 32) based on school scheduling constraints and their availability to initiate the intervention immediately. Two students assigned to the intervention group withdrew before starting the intervention, received no treatment, and did not complete the baseline assessment; therefore, they were excluded from all analyses. The final analytic sample comprised 62 participants, including 30 intervention participants who initiated the intervention and 32 waitlist controls (Figure 1).

No statistically significant between-group differences were observed in age or gender. However, the intervention group demonstrated higher baseline symptom severity than the waitlist control group across several measures, including depressive symptoms, anxiety symptoms, psychological inflexibility, and rumination (Table 1). To assess whether session non-completion was systematic, baseline demographic and clinical characteristics were compared between the 26 participants who completed all eight sessions and the four participants who did not complete the full intervention sequence (Supplementary Table 1). No statistically significant differences were identified for any baseline variable, suggesting that session non-completers did not exhibit a distinct baseline profile. Nevertheless, substantial baseline imbalance between the intervention and waitlist control groups remained a key methodological limitation. Therefore, all outcome analyses were adjusted for baseline scale scores to account for pre-existing symptom differences. The results should be interpreted as adjusted exploratory comparisons rather than unbiased estimates of causal treatment effects.

Table 1 Characteristics of the study sample (n = 62), stratified by intervention group (n = 30) and waitlist control group (n = 32), mean ± SD or n (%).
Characteristic
Control (n = 32)
Intervention (n = 30)
Overall (n = 62)
P value
Age (year)13.81 ± 0.5913.9 ± 0.3113.86 ± 0.470.464
Gender0.362
    Male12 (37.5)8 (26.7)20 (32.3)
    Female20 (62.5)22 (73.3)42 (67.7)
Depression symptoms
    CES-D25.53 ± 5.4734.27 ± 7.4229.76 ± 7.79< 0.001a
    PHQ-99.88 ± 3.3214.8 ± 4.0612.26 ± 4.43< 0.001a
Anxiety symptoms
    GAD-76.91 ± 2.7211.9 ± 4.289.32 ± 4.33< 0.001a
Psychological inflexibility
    AAQ-II22.50 ± 6.8730.37 ± 7.5026.31 ± 8.15< 0.001a
Ruminative thinking
    RRS49.03 ± 10.3758.63 ± 12.4053.68 ± 12.290.002a
Emotion regulation
    ERI41.63 ± 10.5041.43 ± 11.5141.53 ± 10.910.946
Life orientation
    LOT-R11.13 ± 2.2510.00 ± 2.7410.58 ± 2.550.082
Working alliance (WAI-SR)3.56 ± 0.86
    Bond subscale3.96 ± 0.92
    Task subscale3.06 ± 0.87
    Goal subscale3.71 ± 1.09
Engagement
    Number of messages325.60 ± 129.71
    Total use time (second)6080.04 ± 2557.09
Feasibility, engagement, working alliance, and safety-monitoring findings

Working alliance: Participants in the intervention group had a mean total WAI-SR score of 3.56 (SD = 0.86). The mean scores for the Bond, Task, and Goal subscales were 3.96 (SD = 0.92), 3.06 (SD = 0.87), and 3.71 (SD = 1.09), respectively. Overall, these findings suggest that participants reported a moderate perceived working alliance with the chatbot-supported program. However, this measure should not be interpreted as an indicator of treatment efficacy. The distribution of working alliance scores is shown in Figure 4.

Figure 4
Figure 4 Distribution of Working Alliance Inventory-Short Revised total and subscale scores. Box plots show the median (line), interquartile range (box), and values within 1.5 × interquartile range (whiskers). WAI: Working Alliance Inventory.

Engagement: Engagement indicators demonstrated strong short-term participation throughout the 2-week program. Among the 30 participants who initiated the intervention, 26 (86.7%) completed all eight scheduled sessions. Among participants with available system usage data, the mean number of chatbot messages exchanged during the intervention period was 325.60. The mean total usage time was 101.3 minutes (SD = 42.6 minutes), corresponding to approximately 12.7 minutes per scheduled session. Daily engagement trends are illustrated in Figure 5.

Figure 5
Figure 5 Daily user messages and daily chatbot use duration across the intervention period. A: The number of user messages by day in the study (X-axis) and by participant (Y-axis). Color intensity represents the number of messages sent per day; B: The time of use (minutes) by day in the study (X-axis) and by participant (Y-axis). Color intensity represents the time of use per day.

Safety monitoring: All intervention-group conversations were screened for risk-related content, and alerts were generated for nine students. Flagged content included suicidal or self-harm ideation, wishes to die, self-injury, bullying victimization, and family violence. Each alert was communicated to authorized staff through both an in-app message and an SMS notification. School counselors reported that they conducted offline follow-up with these students in accordance with the study protocol. These follow-up procedures occurred outside the platform and were not captured in system logs. Therefore, the logs could not determine the number of students who received follow-up, the specific actions taken, or the time interval between alert generation and follow-up.

This pilot study was not designed to validate the performance of the risk-detection component. No independent reference standard was available, and the May 2025 system version used a binary trigger for dangerous behavior (detected vs not detected) without graded risk stratification. Therefore, no claims are made regarding sensitivity, specificity, false-positive burden, response time, positive predictive value, or missed cases. The 9 generated alerts indicate that the detect-and-notify workflow was operational in practice, but do not provide evidence of validated detection accuracy.

Exploratory clinical outcomes

Adjusted clinical outcome analyses were conducted to identify preliminary outcome patterns relevant to future implementation rather than to evaluate treatment efficacy. Because the intervention group had substantially greater baseline symptom severity than the control group (Table 1), and because multiple outcomes were assessed in a small pilot sample, baseline-adjusted ANCOVA was used as the primary clinical analysis. Results are presented as adjusted between-group mean differences (intervention minus control) with 95%CIs. Time-by-group mixed-model comparisons are reported alongside these estimates to maintain consistency with the original analysis plan. Within-group changes and FDR-corrected P values are provided in Supplementary Table 2. All P values were interpreted descriptively, and no clinical outcome was considered confirmatory.

Anxiety: After adjustment for baseline severity, anxiety did not differ significantly between groups (GAD-7 adjusted mean difference = -0.52, 95%CI: -2.94 to 1.90, P = 0.668; Table 2). The mean change in GAD-7 scores from baseline to post-intervention was -1.47 (SD = 4.08) in the intervention group and 0.75 (SD = 3.86) in the control group. The nominal signal observed in the time-by-group model was not maintained after baseline adjustment and did not remain significant after FDR correction (GAD-7: F = 4.81, P = 0.032; between-group Cohen’s d = -0.56, 95%CI: -1.07 to -0.05; Supplementary Table 2).

Table 2 Comparisons of mental health outcomes between groups across time points.
Measure and group
Baseline, mean ± SD
Post-intervention, mean ± SD
Δmean ± SD
n
Between-group LMM1
ANCOVA2
FCohen’s d3, 95%CIP valueAdjusted mean difference, 95%CIP value
CES-D
Intervention34.27 ± 7.4231.40 ± 9.08-2.87 ± 7.55302.40-0.40, -0.90 to 0.110.127-0.96, -5.96 to 4.040.703
Control25.53 ± 5.4725.94 ± 9.530.41 ± 8.8632
GAD-7
Intervention11.90 ± 4.2810.43 ± 4.78-1.47 ± 4.08304.81-0.56, -1.07 to -0.050.032a-0.52, -2.94 to 1.900.668
Control6.91 ± 2.727.66 ± 4.150.75 ± 3.8632
PHQ-9
Intervention14.80 ± 4.0612.90 ± 5.58-1.90 ± 4.31301.54-0.32, -0.82 to 0.180.220-0.34, -3.46 to 2.780.828
Control9.88 ± 3.329.56 ± 5.66-0.31 ± 5.5032
AAQ-II
Intervention30.37 ± 7.5029.93 ± 8.82-0.43 ± 6.98300.52-0.18, -0.68 to 0.320.4731.11, -3.00 to 5.220.592
Control22.50 ± 6.8723.38 ± 7.820.88 ± 7.7532
RRS
Intervention58.63 ± 12.4054.10 ± 12.51-4.53 ± 11.30302.85-0.44, -0.95 to 0.060.097-1.22, -6.63 to 4.190.654
Control49.03 ± 10.3749.06 ± 11.440.03 ± 9.3432
ERI
Intervention41.43 ± 11.5141.17 ± 11.62-0.27 ± 8.70300.020.04, -0.46 to 0.540.879-0.03, -4.69 to 4.630.990
Control41.63 ± 10.5040.97 ± 10.30-0.66 ± 11.3032
LOT-R
Intervention10.00 ± 2.7411.07 ± 3.211.07 ± 3.51300.930.25, -0.25 to 0.750.3380.14, -1.37 to 1.650.854
Control11.13 ± 2.2511.38 ± 2.760.25 ± 2.9032
LOT-R subscale
    Optimism
        Intervention5.07 ± 1.516.30 ± 2.221.23 ± 2.22305.130.59, 0.08 to 1.090.027a0.71, -0.44 to 1.870.219
        Control6.13 ± 1.566.03 ± 2.10-0.09 ± 2.3132
    Pessimism
        Intervention4.93 ± 2.084.77 ± 2.19-0.17 ± 2.00300.69-0.20, -0.70 to 0.290.409-0.56, -1.59 to 0.460.277
        Control5.00 ± 2.215.34 ± 1.960.34 ± 2.8832

Depressive symptoms: After adjustment for baseline scores, no between-group differences were observed for depressive symptoms measured by the CES-D (adjusted mean difference = -0.96, 95%CI: -5.96 to 4.04, P = 0.703) or PHQ-9 (adjusted mean difference = -0.34, 95%CI: -3.46 to 2.78, P = 0.828; Table 2). No statistically significant time-by-group interaction was identified for CES-D scores (F = 2.40, P = 0.127; between-group Cohen’s d = -0.40, 95%CI: -0.90 to 0.11) or PHQ-9 scores (F = 1.54, P = 0.220; between-group Cohen’s d = -0.32, 95%CI: -0.82 to 0.18). Mean CES-D scores changed by -2.87 (SD = 7.55) in the intervention group and 0.41 (SD = 8.86) in the control group. Mean PHQ-9 scores changed by -1.90 (SD = 4.31) and -0.31 (SD = 5.50), respectively.

Rumination: Rumination showed a nonsignificant trend favoring the intervention group in both analyses (RRS time-by-group model: F = 2.85, P = 0.097; between-group Cohen’s d = -0.44, 95%CI: -0.95 to 0.06; adjusted mean difference = -1.22, 95%CI: -6.63 to 4.19, P = 0.654; Table 2). Mean RRS scores changed by -4.53 (SD = 11.30) in the intervention group and 0.03 (SD = 9.34) in the control group. Although the directional pattern favored the intervention group, this pattern did not provide evidence of a reliable short-term effect on depression or rumination.

Optimism and other secondary outcomes: Optimism showed a nominal time-by-group difference (LOT-R optimism subscale: F = 5.13, P = 0.027; between-group Cohen’s d = 0.59, 95%CI: 0.08-1.09). However, after adjustment for baseline scores, no significant between-group difference was observed for optimism (adjusted mean difference = 0.71, 95%CI: -0.44 to 1.87, P = 0.219; Table 2). Mean optimism scores increased by 1.23 (SD = 2.22) in the intervention group and decreased by 0.09 (SD = 2.31) in the control group. No significant baseline-adjusted between-group differences were observed for psychological inflexibility measured by the AAQ-II (adjusted mean difference = 1.11, 95%CI: -3.00 to 5.22, P = 0.592), emotion regulation measured by the ERI (adjusted mean difference = -0.03, 95%CI: -4.69 to 4.63, P = 0.990), overall life orientation measured by the LOT-R (adjusted mean difference = 0.14, 95%CI: -1.37 to 1.65, P = 0.854), or pessimism (adjusted mean difference = -0.56, 95%CI: -1.59 to 0.46, P = 0.277; Table 2). Overall, after accounting for baseline severity, no outcome demonstrated a statistically significant between-group difference, and no finding remained significant after FDR correction. Therefore, these short-term clinical findings should be considered hypothesis-generating for future randomized trials rather than evidence of treatment benefit.

DISCUSSION
Principal findings

This pilot study primarily provides evidence regarding operational feasibility. A brief GenAI chatbot program could be integrated into the school day, delivered, and monitored within a typical junior high school setting. Among the 30 students who initiated the intervention, 86.7% completed all eight sessions, used Duoduo for approximately 101 minutes over two weeks, and reported a moderate working alliance. These findings support the feasibility of implementation; however, they do not demonstrate that the chatbot reduces depressive or anxiety symptoms.

The clinical findings did not provide evidence of a treatment effect. Anxiety and optimism showed nominal differences in unadjusted comparisons; however, these differences were no longer observed after adjustment for baseline severity. No outcome demonstrated a significant between-group difference after baseline adjustment, and no finding remained significant after FDR correction. Baseline imbalance is a key consideration in interpreting these results. The intervention group had substantially greater symptom severity at baseline than the control group, with CES-D scores of 34.3 vs 25.5, respectively (P < 0.001). In nonrandomized groups selected partly according to baseline status, apparent within-group improvements may reflect regression to the mean rather than intervention effects. The attenuation of the unadjusted signals after baseline adjustment is consistent with this possibility. Given the pragmatic allocation approach, baseline imbalance, small sample size, brief follow-up period, and assessment of multiple outcomes, these findings should be considered hypothesis-generating for future randomized trials rather than evidence of intervention efficacy.

Comparison with previous work

These findings are consistent with previous research in digital mental health, which suggests that youth-focused interventions require more than symptom-focused content to achieve practical effectiveness, with engagement, service context, and implementation fit serving as important determinants[36,37]. Recent studies have similarly highlighted that adolescents’ engagement with digital mental health platforms is influenced by relational experiences, perceived usefulness, and compatibility with the surrounding service environment[10,38]. This study extends previous research by providing a school-based example of how a brief GenAI-supported intervention can be integrated into a supervised routine setting rather than delivered as an open-ended consumer application.

Unlike many consumer-facing chatbots, Duoduo was implemented as a staff-supervised adjunct within a school environment. This distinction is important in digital psychiatry, where factors such as privacy, transparency, safety monitoring, and accountability may be as critical as user experience. In this pilot study, the intervention functioned as an integrated package that combined GenAI-supported conversations, structured ACT-informed sessions, staff review of risk-related content, and school-based follow-up when required.

The ACT framework informed the structure and content of the sessions; however, this pilot study cannot determine which specific components contributed to the observed patterns. Factors including ACT-informed content, multimodal responsiveness, scheduled session delivery, novelty effects, and staff supervision may each have influenced engagement or short-term changes. The nominal improvements observed in anxiety and optimism are consistent with ACT principles, including psychological flexibility, cognitive defusion, acceptance, present-moment awareness, and values-guided action. However, this interpretation remains preliminary and requires further investigation. The multimodal design may have enabled the system to respond to emotional cues that students did not explicitly communicate through text, although this capability also requires additional validation. Among the working-alliance results, the Bond subscale showed the highest score. This finding is relevant to perceived relational connection with a generative system but should not be interpreted as evidence of clinical efficacy.

Implementation implications and cautious clinical interpretation

Overall, the findings reveal that Duoduo is feasible as a school-based digital adjunct. However, the results do not establish Duoduo as an evidence-based treatment or provide an unbiased causal estimate of its clinical effectiveness. Instead, this pilot demonstrates that a brief, structured GenAI program can be integrated into the school schedule and connected to a staff-led review process for risk-related content.

This model may be relevant for school settings where student support needs exceed the capacity of one-to-one counseling services[5]. A supervised digital adjunct may provide an additional access point to support and assist staff in identifying conversations that may require follow-up. A key safeguard of this design is that the chatbot was not intended to operate independently; instead, designated personnel reviewed alerts and determined appropriate follow-up actions.

A further implementation consideration is the workload imposed on school staff by the human-in-the-loop safety model. Because the chatbot was intentionally designed to be non-autonomous, each risk-related alert required timely staff review and, when necessary, school-based follow-up. Therefore, feasibility depends not only on student engagement but also on the capacity of the school system to manage and respond to alerts within the intervention period. In this pilot study, staff time, alert-review turnaround, and follow-up burden were not formally quantified. Future implementation studies should directly assess these parameters, as they are likely to influence whether a supervised GenAI adjunct can be sustainably implemented at scale.

For clinical interpretation, these findings should be considered preliminary. Future studies should determine whether any observed effects are sustained over longer follow-up periods, replicated under randomized allocation, and influenced by factors such as symptom severity, school support capacity, or engagement patterns.

Limitations and future directions

This pilot study has several important limitations. First, participants were assigned pragmatically rather than randomly, resulting in baseline differences between groups. Although baseline adjustment can reduce measured imbalance, it cannot eliminate regression to the mean, residual confounding, or selection effects. Second, the lack of an attention-matched or active-control group prevented separation of chatbot-specific effects from the effects of scheduled sessions, staff contact, counseling-room attendance, and intervention novelty. Supervised contact itself may have contributed to the observed outcomes. Third, multiple outcomes were assessed in a small sample, making nominal P values susceptible to false-positive findings despite FDR correction. Fourth, the sample was recruited from a single school, and follow-up was limited to two weeks, restricting generalizability and likely providing insufficient time to influence more stable psychological constructs such as psychological inflexibility or persistent rumination. Fifth, all outcomes were based on self-report measures, without clinician-rated assessments, diagnostic interviews, or longer-term evaluations of service use.

The intervention combined several components, including ACT content, multimodal responsiveness, scheduled sessions, novelty effects, and staff supervision. Therefore, the observed changes cannot be attributed to any single component, and the specific contributions of ACT content, chatbot delivery, and human oversight cannot be distinguished. Safety monitoring also requires further investigation. Although alerts were integrated into a human-in-the-loop workflow, this pilot did not assess detector sensitivity, specificity, false-positive burden, response time, staff workload, or clinical accuracy. No independent reference standard was available to evaluate missed cases, and offline staff review and follow-up procedures were not recorded within the platform. Future studies should also examine requirements for large-scale implementation, including staff capacity, variability in schools’ ability to manage alerts, and long-term data governance considerations. Finally, acceptability was inferred from completion, engagement, and working-alliance scores rather than assessed using a dedicated satisfaction or usability measure.

Future studies should evaluate this model in randomized, multi-site trials with longer follow-up periods. Additional research should examine safety-alert performance, staff time requirements, data governance procedures, and the consistency of system behavior across diverse student populations. Addressing these issues will be essential for determining whether GenAI-supported tools can progress from small-scale pilot studies to responsibly implemented school-based mental health services.

CONCLUSION

In this single-school, nonrandomized controlled pilot study, a staff-supervised, school-based ACT-informed multimodal GenAI chatbot was feasible to implement among adolescents with elevated depressive symptoms. Feasibility was supported by session completion, engagement, working alliance, and successful operation of the safety monitoring workflow. Acceptability was inferred indirectly from completion rates, engagement indicators, and working-alliance measures rather than assessed using a dedicated satisfaction or usability instrument; therefore, these findings should be interpreted with caution. This study does not establish clinical efficacy or validate the performance of the automated safety-detection system. Symptom outcomes and safety monitoring performance require further evaluation in randomized, multisite trials with longer follow-up periods, direct measures of acceptability, more comprehensive reporting of governance procedures, and clearer separation between system development, implementation, and outcome evaluation.

ACKNOWLEDGEMENTS

We thank the participating students, their guardians, school staff, and school counselors for their support during study delivery and safety follow-up.

References
1.  Wang S, Li Q, Lu J, Ran H, Che Y, Fang D, Liang X, Sun H, Chen L, Peng J, Shi Y, Xiao Y. Treatment Rates for Mental Disorders Among Children and Adolescents: A Systematic Review and Meta-Analysis. JAMA Netw Open. 2023;6:e2338174.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 75]  [Cited by in RCA: 70]  [Article Influence: 23.3]  [Reference Citation Analysis (0)]
2.  Zhou J, Liu Y, Ma J, Feng Z, Hu J, Hu J, Dong B. Prevalence of depressive symptoms among children and adolescents in china: a systematic review and meta-analysis. Child Adolesc Psychiatry Ment Health. 2024;18:150.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 56]  [Reference Citation Analysis (0)]
3.  Luan Y, Wang Y, Chu L, Zhang S, Fan X, Tsin C, Yu Z, Chen K, Wang X, Zhang Y, Wang Y, Gao L, Jiang N, Zhang J, Machavariani E, Qiu Y, Zhong N, Zhao M, Li G, Long J, Jiang F, Lu C. Policies, resources, and interventions for child and adolescent mental health in China: a scoping review. Lancet Reg Health West Pac. 2025;62:101635.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 1]  [Cited by in RCA: 8]  [Article Influence: 8.0]  [Reference Citation Analysis (0)]
4.  McLaughlin KA, Nolen-Hoeksema S. Rumination as a transdiagnostic factor in depression and anxiety. Behav Res Ther. 2011;49:186-193.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 545]  [Cited by in RCA: 502]  [Article Influence: 33.5]  [Reference Citation Analysis (1)]
5.  McPhail L, Thornicroft G, Gronholm PC. Help-seeking processes related to targeted school-based mental health services: systematic review. BMC Public Health. 2024;24:1217.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 19]  [Cited by in RCA: 10]  [Article Influence: 5.0]  [Reference Citation Analysis (0)]
6.  Radez J, Reardon T, Creswell C, Lawrence PJ, Evdoka-Burton G, Waite P. Why do children and adolescents (not) seek and access professional help for their mental health problems? A systematic review of quantitative and qualitative studies. Eur Child Adolesc Psychiatry. 2021;30:183-211.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 855]  [Cited by in RCA: 584]  [Article Influence: 116.8]  [Reference Citation Analysis (5)]
7.  Clement S, Schauman O, Graham T, Maggioni F, Evans-Lacko S, Bezborodovs N, Morgan C, Rüsch N, Brown JS, Thornicroft G. What is the impact of mental health-related stigma on help-seeking? A systematic review of quantitative and qualitative studies. Psychol Med. 2015;45:11-27.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1906]  [Cited by in RCA: 1734]  [Article Influence: 157.6]  [Reference Citation Analysis (3)]
8.  Lehtimaki S, Martic J, Wahl B, Foster KT, Schwalbe N. Evidence on Digital Mental Health Interventions for Adolescents and Young People: Systematic Overview. JMIR Ment Health. 2021;8:e25847.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 176]  [Cited by in RCA: 273]  [Article Influence: 54.6]  [Reference Citation Analysis (0)]
9.  Li H, Zhang R, Lee YC, Kraut RE, Mohr DC. Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being. NPJ Digit Med. 2023;6:236.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 31]  [Cited by in RCA: 249]  [Article Influence: 83.0]  [Reference Citation Analysis (0)]
10.  Escobar-Viera CG, Porta G, Coulter RWS, Martina J, Goldbach J, Rollman BL. A chatbot-delivered intervention for optimizing social media use and reducing perceived isolation among rural-living LGBTQ+ youth: Development, acceptability, usability, satisfaction, and utility. Internet Interv. 2023;34:100668.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 16]  [Reference Citation Analysis (0)]
11.  Zhang Q, Zhang R, Xiong Y, Sui Y, Tong C, Lin FH. Generative AI Mental Health Chatbots as Therapeutic Tools: Systematic Review and Meta-Analysis of Their Role in Reducing Mental Health Issues. J Med Internet Res. 2025;27:e78238.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 49]  [Cited by in RCA: 20]  [Article Influence: 20.0]  [Reference Citation Analysis (0)]
12.  Dray J, Symons D. Review of Innovative Mental Health Support for Children and Young People: Generative AI Co-design Applications and Challenges. Curr Dev Disord Rep. 2025;12:13.  [PubMed]  [DOI]  [Full Text]
13.  Hayes SC. Acceptance and Commitment Therapy, Relational Frame Theory, and the Third Wave of Behavioral and Cognitive Therapies - Republished Article. Behav Ther. 2016;47:869-885.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 217]  [Cited by in RCA: 465]  [Article Influence: 46.5]  [Reference Citation Analysis (1)]
14.  Ma J, Ji L, Lu G. Adolescents' experiences of acceptance and commitment therapy for depression: An interpretative phenomenological analysis of good-outcome cases. Front Psychol. 2023;14:1050227.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in RCA: 7]  [Reference Citation Analysis (0)]
15.  López-Pinar C, Lara-Merín L, Macías J. Process of change and efficacy of acceptance and commitment therapy (ACT) for anxiety and depression symptoms in adolescents: A meta-analysis of randomized controlled trials. J Affect Disord. 2025;368:633-644.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 10]  [Cited by in RCA: 12]  [Article Influence: 12.0]  [Reference Citation Analysis (0)]
16.  Fitzpatrick KK, Darcy A, Vierhile M. Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial. JMIR Ment Health. 2017;4:e19.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 766]  [Cited by in RCA: 805]  [Article Influence: 89.4]  [Reference Citation Analysis (6)]
17.  Hawke LD, Hou J, Nguyen ATP, Phi T, Gibson J, Ritchie B, Strudwick G, Rodak T, Gallagher L. Digital Conversational Agents for the Mental Health of Treatment-Seeking Youth: Scoping Review. JMIR Ment Health. 2025;12:e77098.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 4]  [Reference Citation Analysis (0)]
18.  Feng X, Tian L, Ho GWK, Yorke J, Hui V. The Effectiveness of AI Chatbots in Alleviating Mental Distress and Promoting Health Behaviors Among Adolescents and Young Adults: Systematic Review and Meta-Analysis. J Med Internet Res. 2025;27:e79850.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 15]  [Reference Citation Analysis (0)]
19.  Udahemuka G, Djouani K, Kurien AM. Multimodal Emotion Recognition Using Visual, Vocal and Physiological Signals: A Review. Appl Sci. 2024;14:8071.  [PubMed]  [DOI]  [Full Text]
20.  Al-Saadawi HFT, Das B, Das R. A systematic review of trimodal affective computing approaches: Text, audio, and visual integration in emotion recognition and sentiment analysis. Expert Syst Appl. 2024;255:124852.  [PubMed]  [DOI]  [Full Text]
21.  Wang X, Zhou Y, Zhou G. The Application and Ethical Implication of Generative AI in Mental Health: Systematic Review. JMIR Ment Health. 2025;12:e70610.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 76]  [Cited by in RCA: 21]  [Article Influence: 21.0]  [Reference Citation Analysis (0)]
22.  Cao XJ, Liu XQ. Artificial intelligence-assisted psychosis risk screening in adolescents: Practices and challenges. World J Psychiatry. 2022;12:1287-1297.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in CrossRef: 59]  [Cited by in RCA: 42]  [Article Influence: 10.5]  [Reference Citation Analysis (1)]
23.  Munder T, Wilmers F, Leonhart R, Linster HW, Barth J. Working Alliance Inventory-Short Revised (WAI-SR): psychometric properties in outpatients and inpatients. Clin Psychol Psychother. 2010;17:231-239.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 442]  [Cited by in RCA: 268]  [Article Influence: 16.8]  [Reference Citation Analysis (0)]
24.  Hatcher RL, Gillaspy JA. Development and validation of a revised short version of the working alliance inventory. Psychother Res. 2006;16:12-25.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 604]  [Cited by in RCA: 745]  [Article Influence: 37.3]  [Reference Citation Analysis (0)]
25.  Radloff LS. The CES-D Scale: A Self-Report Depression Scale for Research in the General Population: A Self-Report Depression Scale for Research in the General Population. Appl Psych Meas. 1977;1:385-401.  [PubMed]  [DOI]  [Full Text]
26.  Yang W, Xiong G, Garrido LE, Zhang JX, Wang MC, Wang C. Factor structure and criterion validity across the full scale and ten short forms of the CES-D among Chinese adolescents. Psychol Assess. 2018;30:1186-1198.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 23]  [Cited by in RCA: 70]  [Article Influence: 8.8]  [Reference Citation Analysis (0)]
27.  Kroenke K, Spitzer RL, Williams JB. The PHQ-9: validity of a brief depression severity measure. J Gen Intern Med. 2001;16:606-613.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 37319]  [Cited by in RCA: 33160]  [Article Influence: 1326.4]  [Reference Citation Analysis (9)]
28.  Tsai FJ, Huang YH, Liu HC, Huang KY, Huang YH, Liu SI. Patient health questionnaire for school-based depression screening among Chinese adolescents. Pediatrics. 2014;133:e402-e409.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 68]  [Cited by in RCA: 123]  [Article Influence: 10.3]  [Reference Citation Analysis (0)]
29.  Spitzer RL, Kroenke K, Williams JB, Löwe B. A brief measure for assessing generalized anxiety disorder: the GAD-7. Arch Intern Med. 2006;166:1092-1097.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 26995]  [Cited by in RCA: 22510]  [Article Influence: 1125.5]  [Reference Citation Analysis (23)]
30.  Sun J, Liang K, Chi X, Chen S. Psychometric Properties of the Generalized Anxiety Disorder Scale-7 Item (GAD-7) in a Large Sample of Chinese Adolescents. Healthcare (Basel). 2021;9:1709.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 225]  [Cited by in RCA: 206]  [Article Influence: 41.2]  [Reference Citation Analysis (1)]
31.  Treynor W, Gonzalez R, Nolen-Hoeksema S. Rumination reconsidered: A psychometric analysis. Cognit Ther Res. 2003;27:247-259.  [PubMed]  [DOI]  [Full Text]
32.  Bond FW, Hayes SC, Baer RA, Carpenter KM, Guenole N, Orcutt HK, Waltz T, Zettle RD. Preliminary psychometric properties of the Acceptance and Action Questionnaire-II: a revised measure of psychological inflexibility and experiential avoidance. Behav Ther. 2011;42:676-688.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 1944]  [Cited by in RCA: 2246]  [Article Influence: 149.7]  [Reference Citation Analysis (1)]
33.  Roth G, Assor A, Niemiec CP, Deci EL, Ryan RM. The emotional and academic consequences of parental conditional regard: comparing conditional positive regard, conditional negative regard, and autonomy support as parenting practices. Dev Psychol. 2009;45:1119-1142.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 232]  [Cited by in RCA: 149]  [Article Influence: 8.8]  [Reference Citation Analysis (0)]
34.  Glaesmer H, Rief W, Martin A, Mewes R, Brähler E, Zenger M, Hinz A. Psychometric properties and population-based norms of the Life Orientation Test Revised (LOT-R). Br J Health Psychol. 2012;17:432-445.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Cited by in Crossref: 266]  [Cited by in RCA: 213]  [Article Influence: 15.2]  [Reference Citation Analysis (0)]
35.  Scheier MF, Carver CS, Bridges MW. Distinguishing optimism from neuroticism (and trait anxiety, self-mastery, and self-esteem): a reevaluation of the Life Orientation Test. J Pers Soc Psychol. 1994;67:1063-1078.  [PubMed]  [DOI]  [Full Text]
36.  Cohen KA, Schleider JL. Adolescent dropout from brief digital mental health interventions within and beyond randomized trials. Internet Interv. 2022;27:100496.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 5]  [Cited by in RCA: 42]  [Article Influence: 10.5]  [Reference Citation Analysis (0)]
37.  Wu Y, Fenfen E, Wang Y, Xu M, Liu S, Zhou L, Song G, Shang X, Yang C, Yang K, Li X. Efficacy of internet-based cognitive-behavioral therapy for depression in adolescents: A systematic review and meta-analysis. Internet Interv. 2023;34:100673.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in RCA: 23]  [Reference Citation Analysis (0)]
38.  Malouin-Lachance A, Capolupo J, Laplante C, Hudon A. Does the Digital Therapeutic Alliance Exist? Integrative Review. JMIR Ment Health. 2025;12:e69294.  [RCA]  [PubMed]  [DOI]  [Full Text]  [Full Text (PDF)]  [Cited by in Crossref: 109]  [Cited by in RCA: 42]  [Article Influence: 42.0]  [Reference Citation Analysis (0)]
Footnotes

Peer review: Externally peer reviewed.

Peer-review model: Single blind

Specialty type: Psychology

Country of origin: China

Peer-review report’s classification

Scientific quality: Grade C, Grade C

Novelty: Grade A, Grade C

Creativity or innovation: Grade B, Grade B

Scientific significance: Grade A, Grade C

P-Reviewer: Barrios-Martínez DD, Academic Fellow, Adjunct Professor, Affiliate Associate Professor, Chief, Chief Physician, Director, Full Professor, Manager, MD, Principal Investigator, Professor, Research Dean, Researcher, Senior Scientist, Colombia; Jamaluddin J, Assistant Professor, MD, Malaysia S-Editor: Lin C L-Editor: A P-Editor: Zhao S

Write to the Help Desk