This study aimed to identify factors related to happiness among Korean ninth-grade students using machine learning. Data were drawn from the 16th wave (2023) of the Panel Study on Korean Children (PSKC), conducted by the Korea Institute of Child Care and Education, and the final sample included 1,154 ninth-grade students. Happiness was used as the outcome variable, while sociodemographic, psychological, academic, relational, and lifestyle characteristics were included as predictors. Multiple linear regression, Elastic Net, Random Forest, and XGBoost models were compared. Feature importance, SHapley Additive exPlanations (SHAP), and Random Forest partial dependence plots (PDPs) were used to examine predictor importance, direction of predictive contribution, and nonlinear prediction patterns. Model performance rankings differed between cross-validation and final validation, with the models demonstrating broadly comparable predictive performance. In the tree-based machine learning models—Random Forest and XGBoost—self-esteem was the most important predictor, followed by depression, peer attachment, ego-resilience, academic stress, paternal and maternal attachment, self-regulated learning, smartphone overdependence, and subjective socioeconomic status. SHAP analyses showed broadly consistent directions of predictive contribution across the two models, while the Random Forest PDPs indicated nonlinear prediction patterns for several major variables. These findings demonstrate that happiness among Korean ninth-grade students is associated with diverse psychological, relational, and academic factors and provide foundational evidence for understanding adolescent happiness.
본 연구는 중학교 3학년 청소년의 행복감 관련 요인을 머신러닝을 활용하여 파악하기 위해 육아정책연구소의 2023년 제16차 한국아동패널 Ⅱ 자료를 이용하였으며, 분석대상은 중학교 3학년 학생 1,154명이었다. 행복감을 결과변수로 하고, 청소년의 인구사회학적·심리적·학업적·관계적 및 생활습관 특성을 예측변수로 구성하였다. 분석에서는 평균예측모형을 기준으로 다중선형회귀, 엘라스틱넷, 랜덤포레스트 및 XGBoost의 예측 성능을 비교하였다. 또한 변수 중요도와 SHAP 분석을 통해 주요 변수의 상대적 중요도와 예측 기여 방향을 확인하고, 랜덤포레스트의 부분의존도 분석을 통해 비선형적 예측 양상을 탐색하였다. 분석 결과, 교차검증과 최종 검증에서 모형 간 성능 순위는 일관되지 않았으며 전반적으로 유사한 수준을 보였다. 트리 기반 머신러닝 모형(랜덤포레스트 및 XGBoost)에서 자아존중감이 가장 중요한 변수로 나타났으며, 우울, 또래애착, 자아탄력성, 학업스트레스, 부모애착, 자기조절학습, 스마트폰 과의존 및 주관적 사회경제적 지위도 주요 관련 요인으로 확인되었다. SHAP 분석에서 주요 변수의 기여 방향은 두 모형 간 대체로 일관되었으며, 부분의존도 분석에서는 일부 주요 변수와 행복감 간의 비선형적 관계 양상이 나타났다. 이러한 결과는 중학교 3학년 청소년의 행복감이 다양한 심리적·관계적·학업적 요인과 관련되어 있음을 보여주며, 청소년 행복감 이해를 위한 기초자료를 제공한다.
This study aimed to identify latent subtypes of happiness trajectories during early adolescence and to explore key predictors of trajectory membership using machine learning approaches. Data were drawn from Waves 12 to 15 of the Panel Study on Korean Children (PSKC), corresponding to grades 5 through 8. Latent Class Growth Analysis (LCGA) was conducted to identify distinct trajectories of happiness over time. Baseline individual, family, peer, environmental, and future-oriented variables were then used to develop machine learning models predicting trajectory membership. Multiple models were compared, and SHAP (SHapley Additive exPlanations) analysis was applied to the best-performing model to examine the relative importance of predictors. Three distinct happiness trajectories were identified: a high-level slow-decline group, a mid-level average-decline group, and a low-level rapid-decline group. Among the machine learning models, the Random Forest demonstrated the most stable predictive performance. SHAP analysis indicated that multidimensional factors contributed to classifying different happiness trajectories. The findings suggest that changes in overall happiness during early adolescence can be characterized by distinct latent trajectories with different initial levels and rates of decline. Furthermore, the use of machine learning models allowed for a more flexible prediction of happiness trajectories by accounting for the combined influence of individual, family, peer, environmental, and expectations and perceptions about the future.
본 연구는 초기 청소년의 전반적 행복감 변화 궤적의 하위 유형을 규명하고, 머신러닝을 활용하여 행복감 궤적 유형을 예측하는 주요 변인을 탐색하는 것을 목적으로 하였다. 한국아동패널조사(KCPS) 12-15차 자료(초5-중2)를 활용하여 잠재계층성장분석을 실시하여 행복감 변화 궤적 유형을 도출하였다. 기저선 시점의 개인, 가족, 또래, 환경 및 미래인식 관련 변인을 투입하여 행복감 궤적 집단을 예측하는 머신러닝 모델들을 구축·비교하였으며, 최적 모델에 대해서 SHAP 분석을 실시하였다. 초기 청소년기의 행복감 변화 궤적은 고수준-완만감소형, 중간수준-평균감소형, 저수준-빠른감소형의 세집단으로 분류되었다. 머신러닝 모델 중에서는 랜덤포레스트가 가장 안정적인 예측성능을 보였으며, 상대적 중요도 분석 결과, 다차원적 요인이 행복감 궤적 분류에 기여하는 것으로 나타났다. 본 연구는 초기 청소년기의 전반적 행복감 변화가 서로 다른 수준과 감소속도를 지닌 잠재집단으로 구분됨을 확인하였으며, 머신러닝 모델을 통해 개인, 가족, 또래, 환경 및 미래 관련 변인이 복합적으로 작용하는 행복감 변화 궤적을 유연하게 예측할 수 있음을 보여주었다.
This study aims to develop a severity-adjusted length of stay predictive model according to comorbidity index by using machine learning and propose a algorithm of severity-adjusted length of stay (LOS) predictive model. The dataset was taken from Korea Centers for Disease Control and Prevention database of the hospital discharge survey from 2006 to 2015 and the severity-adjusted length of stay predictive model was developed for the nervous system patients to need a urgent management for length of stay. when it comes to the severity-adjusted length of stay predictive model about nervous system discharging patients, three tools were used for the severity-adjustment of comorbidity: the CCI, the ECI, and the CCS. The models using Regression, Decision Tree, Random Forest, Support Vector Regression, Neural Network as a Machine learning analysis methods were developed and then evaluate. As a result, Severity-adjusted predictive model using CCS as the severity-adjustment of comorbidity and Neural Network method has the highest R-square and has the most excellent prediction capability. In conclusion, there is a need to develop a severity-adjusted predictive model using CCS as the severity-adjustment of comorbidity and make use of severity-adjusted predictive model to has high prediction capability by using various machine-learning analytics.
본 연구는 머신러닝을 이용하여 동반상병 보정 방법에 따른 중증도 보정 재원일수 예측 모형을 개발하고 이를 평가하여 중증도 보정 재원일수 예측 모형 개발의 알고리즘을 제시하기 위해 수행되었다. 본 연구를 위해 2006년부터 2015년까지 10년간의 질병관리본부 퇴원손상심층조사 자료를 수집하였으며, 재원일수 관리가 시급한 신경계통의 질환을 대상으로 중증도 보정 재원일수 예측 모형을 개발하였다. 신경계통의 질환 퇴원환자의 중증도 보정 재원일수 예측 모형 개발 시 동반상병 보정 방법은 CCI, ECI, CCS 진단군 분류 기준 등 3가지, 머신러닝 분석기법으로는 회귀분석, 의사결정나무, 랜덤 포레스트, 서포트 백터 회귀분석, 신경망 등 5가지를 적용하여 모형을 개발하고 개발된 모형을 평가하였다. 모형 평가 결과 CCS 진단군 분류 기준 동반상병 보정 방법 및 신경망을 이용하여 개발한 중증도 보정 예측 모형의 모형 설명력(R-square)이 가장 높았으며, 모형의 예측력이 가장 우수한 것으로 나타났다. 따라서 중증도 보정 재원일수 예측 모형 개발 시 CCS 진단군 분류 변수를 이용한 동반상병 보정 방법을 이용하여 중증도 보정 예측 모형을 개발하는 것이 필요하며, 머신러닝의 다양한 분석 기법 등을 이용하여 예측력 높은 중증도 보정 예측 모형을 개발하여 재원일수 변이요인 파악 등 재원일수 관리를 위해 활용하는 것이 필요하다.