IAES
Inter
national
J
our
nal
of
Articial
Intelligence
(IJ-AI)
V
ol.
15,
No.
4,
August
2026,
pp.
3888
∼
3902
ISSN:
2252-8938,
DOI:
10.11591/ijai.v15.i4.pp3888-3902
❒
3888
Pr
edicti
v
e
model
based
on
machine
lear
ning
to
identify
sleep-r
elated
health
pr
oblems
Laberiano
Andrade-Ar
enas
1
,
Inoc
Rubio
P
aucar
2
,
Mar
garita
Giraldo
Retuerto
1
,
Cesar
Y
actay
o-Arias
3
1
F
aculty
of
Science
and
Engineering,
Uni
v
ersidad
de
Ciencias
y
Humanidades,
Lima,
Per
´
u
2
F
aculty
of
Engineering
and
Business,
Uni
v
ersidad
Pri
v
ada
Norbert
W
iener
,
Lima,
Per
´
u
3
Department
of
General
Studies,
Uni
v
ersidad
Continental,
Lima,
Per
´
u
Article
Inf
o
Article
history:
Recei
v
ed
Jul
19,
2025
Re
vised
Jul
13,
2026
Accepted
Jul
21,
2026
K
eyw
ords:
Extreme
gradient
boosting
Kno
wledge
disco
v
ery
in
databases
methodology
Machine
learning
Predicti
v
e
model
Public
health
monitoring
Sleep
quality
ABSTRA
CT
Sleep
quality
has
become
a
gro
wing
public
health
issue
w
orldwide,
mainly
due
to
a
lack
of
a
w
areness
about
its
long-term
consequences.
Despite
e
xisting
strate
gies
to
address
this
problem,
there
remains
a
need
for
more
ef
fecti
v
e
approaches.
In
this
study
,
an
early
detection
model
for
sleep
disorders
w
as
implemented
using
the
e
xtreme
gradient
boosting
(XGBoost)
algorithm,
follo
wing
the
kno
wledge
disco
v
ery
in
databases
(KDD)
methodology
,
which
includes
the
phases
of
selection,
preprocessing,
transformation,
data
mining,
and
interpretation.
A
dataset
e
xtracted
from
the
Kaggle
platform
in
CSV
format
w
as
used,
consisting
of
374
records.
W
ith
an
o
v
erall
accurac
y
of
91.5%,
a
recall
of
100%
for
the
insomnia
class,
and
a
precision
of
100%
for
sleep
apnea,
the
proposed
model
demonstrated
e
xceptional
performance.
It
also
recei
v
ed
an
area
under
the
curv
e
(A
UC)
of
0.909
and
an
a
v
erage
F1-score
of
0.913.
W
ith
a
mean
accurac
y
of
91%,
a
95%
condence
interv
al
(0.89–0.94),
and
a
p-v
alue
of
0.0012,
cross-v
alidation
conrmed
its
rob
ustness
and
sho
wed
a
statistically
signicant
change
from
the
baseline
model.
The
error
rates
remained
withi
n
clinically
acceptable
ranges,
conrming
its
applicability
as
a
diagnostic
support
tool.
Ov
erall,
the
results
demonstrate
the
ef
fecti
v
eness
of
the
model
in
identifying
patterns
related
to
sleep
disorders.
This
is
an
open
access
article
under
the
CC
BY
-SA
license
.
Corresponding
A
uthor:
Cesar
Y
actayo-Arias
Department
of
General
Studies,
Uni
v
ersidad
Continental
Lima,
Per
´
u
Email:
c
yactayo@continental.edu.pe
1.
INTR
ODUCTION
The
scientic
community
is
currently
becoming
more
interested
in
learning
more
about
human
sleep
indicators.
Numerous
studies
ha
v
e
demonstrated
that
sleep
quality
signicantly
af
fects
one’
s
ph
ysical
and
mental
health.
Lifestyle
choices,
eating
habits,
stress
le
v
els,
and
particularly
the
frequenc
y
and
intensity
of
personal
ph
ysical
acti
vity
are
among
the
primary
determinants
of
this
quality
[1],
[2].
In
response
to
this
issue,
data
science
and
machine
learning
(ML)
ha
v
e
emer
ged
as
essential
tools
for
addressing
the
la
r
ge-
scale
analysis
of
biometric
and
beha
vioral
data
related
to
sleep.
These
technologies
also
enable
the
identication
of
comple
x
patterns
and
the
e
xtraction
of
k
e
y
v
ariables
that
f
acilitate
informed
decision-making
in
the
elds
of
public
health
and
the
pre
v
ention
of
sleep
disorders.
The
a
v
ailability
of
v
aluable
data
on
sleep
habits,
the
calculation
of
rele
v
ant
ph
ysiological
indicators,
and
the
assessment
of
lifestyle
choices
pro
vide
an
unprecedented
opportunity
to
in
v
estig
ate
risk
f
actors
that
af
fect
the
o
v
erall
health
of
the
population.
J
ournal
homepage:
http://ijai.iaescor
e
.com
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
3889
Although
there
is
e
xtensi
v
e
e
vidence
linking
sleep
to
v
arious
aspects
of
health,
there
are
stil
l
challenges
in
accurately
predicting
the
potential
ne
g
ati
v
e
health
consequences
caused
by
poor
sleep
patterns.
Moreo
v
er
,
man
y
studies
rely
on
traditional
statistical
analyses,
whose
main
limitation
lies
in
their
inability
to
capture
nonlinear
relationships
between
the
studied
v
ariables
[3],
[4].
F
or
this
reason,
there
is
a
gro
wing
interest
in
complementing
these
approaches
with
ML
techniques,
which
enable
the
modeling
of
comple
x
patterns
and
the
deli
v
ery
of
more
precise
and
in-depth
kno
wledge.
Sleep
is
a
vital
ph
ysiological
function
for
human
beings,
and
its
sustained
depri
v
ation
or
disruption
can
signicantly
compromise
functional
performance
in
daily
acti
vities
and
promote
the
onset
of
v
arious
non-communicable
chronic
diseases
[5].
In
this
re
g
ard,
scientic
e
vidence
has
sho
wn
that
sleep
disorders
are
closely
associated
with
an
increased
risk
of
de
v
eloping
h
ypertension,
t
yp
e
2
diabetes
mellitus,
obesity
,
as
well
as
mental
health
disorders,
including
anxiety
and
depression.
Despite
thi
s,
man
y
indi
v
i
duals
prioritize
their
daily
responsibilities
without
considering
that
chronic
sleep
restriction
represents
a
signicant
risk
f
actor
for
their
o
v
erall
well-being.
This
ne
glect
has
contrib
uted
to
a
high
pre
v
alence
of
chronic
stress
conditions
and
sleep-related
pathologies
[6].
In
light
of
this,
people
must
become
a
w
are
of
the
k
e
y
aspects
of
achie
ving
proper
sleep
at
appropriate
times
in
order
to
maintain
a
health
y
life.
In
this
conte
xt,
the
application
of
ML
techniques,
framed
within
the
kno
wledge
disco
v
ery
in
dat
abases
(KDD)
met
ho
dol
ogy
,
enables
the
e
xploration
and
analysis
of
lar
ge
v
olumes
of
data
with
the
aim
of
identifying
non-tri
vial
patterns
and
de
v
eloping
highly
accurate
predicti
v
e
models.
This
approach
not
only
strengthens
research
in
public
health
and
beha
vioral
sciences
b
ut
also
pro
vides
adv
anced
analytical
tools
for
pre
v
ention,
early
diagnosis,
and
the
personalization
of
health
interv
entions
[7],
[8].
It
is
important
to
consider
that
ML
is
a
k
e
y
f
actor
in
generating
kno
wledge
that
adds
v
alue
in
addressing
a
public
health
issue
[9],
[10].
The
present
research
is
justied
by
its
capacity
to
generate
applied
kno
wledge
in
the
eld
of
pre
v
enti
v
e
health
through
the
design
and
implementation
of
an
algorithmic
model
capable
of
predicting
health
risks
based
on
the
analysis
of
v
ariables
related
to
sleep
quality
and
lifestyle.
This
contrib
ution
not
only
e
xpands
the
s
cientic
e
vidence
base
b
ut
also
pro
vides
rele
v
ant
inputs
for
decision-making
by
healthcare
professionals,
public
polic
y
mak
ers
,
and
de
v
elopers
of
technologies
aimed
at
promoting
population
well-being.
The
objecti
v
e
of
this
research
is
to
de
v
elop
a
predicti
v
e
model
using
ML
techniques
specical
ly
the
e
xtreme
gradient
boosting
(XGBoost)
algorithm
that
enables
the
identication
and
pre
diction
of
potential
impacts
on
ph
ysical
and
mental
health
based
on
the
analysis
of
sleep
patterns.
2.
LITERA
TURE
REVIEW
This
section
focuses
on
related
w
orks
by
v
arious
authors
on
the
topic,
thus
pro
viding
a
holi
stic
o
v
ervie
w
for
the
in
v
estig
ation.
Lik
e
wi
se,
the
theoretical
foundations
s
upp
or
t
the
identication
and
justication
of
the
v
ariables
to
be
studied.
This
comprehensi
v
e
approach
ensures
a
solid
frame
w
ork
for
the
current
study
.
2.1.
Related
w
orks
The
article
e
v
aluates
v
arious
ML
models
for
predicting
sleep
disorders,
considering
their
si
gnicant
impact
on
public
health.
The
study
utilizes
a
dataset
of
400
indi
vidual
records
classied
by
sleep
disorder
status
(none,
insomnia,
and
sleep
apnea)
and
incorporates
demographic,
lifestyle,
and
health
metrics.
Dif
ferent
ML
models,
including
logistic
re
gression,
decision
trees,
ensemble
methods
(lik
e
random
forest
(RF)
and
gradient
boosting),
support
v
ector
machines,
and
neural
netw
orks
are
assessed
for
their
diagnostic
accurac
y
.
K
e
y
ndings
indicate
that
ensemble
methods,
particularly
RF
and
e
xtreme
gradient
boosting
classier
(XGBClassier),
outperform
other
models,
achie
ving
high
accurac
y
and
F1-scores
(up
to
0.93).
The
results
highlight
the
importance
of
adv
anced
ensemble
techniques
in
accurately
diagnosing
sleep
disorders,
suggesting
their
inte
gration
into
clinical
practices
for
impro
v
ed
outcomes.
The
research
methodology
includes
data
collection
and
preprocessing,
model
de
v
elopment,
trai
ning,
and
e
v
aluation
using
metrics
lik
e
accurac
y
,
precision,
recall,
and
F1-score.
The
study
emphasizes
the
comple
xity
of
sleep
disorders
and
the
potential
of
ML
to
enhance
diagnostic
tools
in
sleep
medicine,
adv
ocating
for
future
research
into
optimizing
these
models
for
real-w
orld
applications
[11],
[12].
Sleep
problems
pose
an
important
challenge
for
human
health.
In
this
study
,
a
cross
-v
alidated
XGBoost–bidirectional
long
short-term
memory
(BiLSTM)
model
w
as
implemented
using
75
v
ariables
and
a
dataset
of
10,765
aging-related
records.
The
proposed
model
reached
97%
accurac
y
and
demonstrated
better
performance
than
traditional
long
short-term
memory
(LSTM)
models
and
other
machine-learning
approaches.
The
results
demonstrate
its
ef
fecti
v
eness
in
the
diagnosis
and
treatment
of
sleep
apnea
by
identifying
risk
Pr
edictive
model
based
on
mac
hine
learning
to
identify
sleep-r
elated
...
(Laberiano
Andr
ade-Ar
enas)
Evaluation Warning : The document was created with Spire.PDF for Python.
3890
❒
ISSN:
2252-8938
f
actors.
In
pre
vious
studies,
the
XGBoost
algorithm
w
as
also
emplo
yed,
applying
training
strate
gies
focused
on
the
most
rele
v
ant
features
of
the
dataset.
As
in
the
pre
vious
study
,
k-fold
cross-v
alidation
w
as
implemented
to
optimize
model
performance.
The
results
were
obt
ained
by
retraining
on
subsets
of
the
training
set
using
test
data,
achie
ving
an
accurac
y
of
90.6%
in
predicting
the
dif
ferent
stages
of
sleep
[13],
[14].
On
the
other
hand,
some
authors
in
v
estig
ated
obstructi
v
e
sleep
apnea
(OSA)
in
the
pediatric
population,
linking
it
to
de
v
elopmental
aspects
such
as
ph
ysical
gro
wth,
cardio
v
ascular
function,
and
cognition
through
ML
techniques.
Using
a
sample
of
3,139
children,
the
XGBoost
algorithm
w
as
applied
after
randomly
splitting
the
data
into
training
and
testing
sets.
In
the
testing
set,
the
model
achie
v
ed
area
under
the
curv
e
(A
UC)
of
0.95,
0.88,
and
0.88
for
the
classication
of
mild,
moderate,
and
se
v
ere
OSA,
respecti
v
ely
,
with
accuracies
of
90.45%,
85.67%,
and
89.81%,
outperforming
re
gression-based
models.
Similarly
,
complementary
approaches
ha
v
e
been
e
xplored
using
techniques
such
as
entrop
y
and
linear
feature
analysis
for
the
selection
of
rele
v
ant
v
ariables.
In
this
conte
xt,
a
model
based
on
the
XGBoost
algorithm
w
as
trained
using
the
Sleep
Heart
Health
Study
2
dataset.
Through
feature
selection
strate
gies,
the
most
signicant
attrib
utes
for
diagnosis
were
identied.
When
e
v
aluating
the
model
with
a
v
alidation
subset
composed
of
422
participants
not
included
in
the
trai
n
i
ng
phase,
an
accurac
y
of
68.83%,
a
precision
of
71.61%,
and
a
recall
of
67.84%
were
achie
v
ed,
demonstrating
acceptable
performance
in
scenarios
in
v
olving
pre
viously
unseen
data,
as
reported
by
[15],
[16].
In
contrast
to
other
studies
[17],
[18]
conducted
a
longitudinal
analysis
on
239
patients
with
insomnia
using
the
XGBoost
algorithm,
which
achie
v
ed
a
multiclass
A
UC
of
80%
and
a
binary
A
UC
of
83%,
with
adequate
sensiti
vity
and
specicity
,
f
acilitating
the
identication
of
patients
with
lo
w
response
to
con
v
entional
depression
treatments.
In
line
with
pre
vious
studies
on
the
relationship
between
sleep-w
ak
e
c
ycle
disorders
and
major
depressi
v
e
disorders
in
adolescents,
v
arious
ML
models
were
e
v
aluated,
with
XGBoost
standing
out
by
achie
ving
an
A
UC
of
85%
in
predicting
the
risk
of
depression.
Mul
tiple
in
v
estig
ations
ha
v
e
demonstrated
that
the
XGBoost
algorithm
is
highly
ef
fecti
v
e
in
anticipating
the
onset
of
depressi
v
e
disorders
associated
with
sleep-w
ak
e
rh
ythm
disturbances,
thus
supporting
the
de
v
elopment
of
personalized
treatments
and
more
precise
pre
v
enti
v
e
strate
gies.
Lik
e
wise,
conditions
such
as
OSA,
the
comorbidi
ty
of
insomnia
and
sleep
apnea
(COMISA),
and
insomnia
a
s
an
independent
disorder
represent
signicant
manifestations
of
sleep
disorders.
In
order
to
predict
the
risk
of
these
conditions,
XGBoost
w
as
implemented
on
clinical
data
collected
from
medical
centers,
selecting
nine
rele
v
ant
v
ariables
through
the
Shaple
y
additi
v
e
e
xplanations
(SHAP)
e
xplainability
method.
These
features
were
used
to
construct
the
schedule,
light,
eat
and
drink,
en
vironment,
ph
ysiology
,
and
stress
(SLEEPS)
model,
which
demonstrated
outstanding
predicti
v
e
capacity
,
with
A
UC
v
alues
e
xceeding
89%
across
all
three
analyzed
disorders.
As
a
result,
SLEEPS
is
positioned
as
an
inno
v
ati
v
e
diagnostic
tool
that
enables
early
and
accessible
detection
of
sleep
disorders
without
relying
on
in
v
asi
v
e
methods
such
as
polysomnograph
y
.
In
the
present
s
tudy
,
three
predicti
v
e
models
were
de
v
eloped
using
the
Python
programming
language:
logistic
re
gression,
XGBClassier
,
and
support
v
ector
classier
(SVC).
After
e
v
aluating
their
performance,
the
XGBClassier
model
w
as
i
dentied
as
the
most
ef
fecti
v
e.
The
results
sho
wed
that
occupation
is
the
main
f
actor
associated
with
insomnia,
while
body
mass
inde
x
(BMI)
cate
gory
signicantly
inuences
the
occurrence
of
sleep
apnea.
Additionally
,
BMI,
stress
le
v
el,
blood
pressure,
and
sleep
duration
were
identied
as
k
e
y
v
ariables
af
fecting
sleep
quality
.
These
ndings
suggest
that
interv
entions
aimed
at
controlling
BMI
and
blood
pressure
could
signicantly
contrib
ute
to
impro
ving
nighttime
rest
[19],
[20].
On
the
other
hand,
a
study
focused
on
estimating
the
risk
of
insomnia
in
breast
cancer
patients
used
data
obtained
through
surv
e
ys.
In
the
predictions
performed,
logistic
re
gression
models
with
L2
re
gularization
and
the
XGBoost
algorithm
achie
v
ed
predicti
v
e
accuracies
of
71.5%
and
70.6%,
respecti
v
ely
,
with
A
UC
v
alues
of
0.76
and
0.75.
Additionally
,
population
subgroups
at
high
risk
of
insom
n
i
a
were
found
by
using
the
RuleFit
algorithm.
The
ndings
demonstrated
that
among
breast
cancer
survi
v
ors,
cancer
-related
f
atigue
is
a
signicant
predictor
of
the
onset
of
sleeplessness.
In
contrast
to
the
study
that
e
xamined
insomnia
in
patients
with
breast
cancer
,
the
current
study
used
the
XGBoost
algorithm
to
nd
risk
f
actors
link
ed
to
sleep
disturbances.
A
total
of
7,929
patients
who
satised
the
inclusion
criteria
were
included
in
the
sample;
their
mean
age
w
as
49.2
years
(standard
de
viation
(SD)
=
18.4),
with
4,055
(51%)
for
w
omen
and
3,874
(49%)
for
men.
The
ethnic
composition
w
as
distrib
uted
as
follo
ws:
2,885
(36%)
white
indi
viduals,
2,144
(27%)
African
Americans,
1,639
(21%)
Hispanics,
and
1,261
(16%)
from
other
ethnic
groups.
Of
the
684
v
ariables
analyzed,
64
were
statistically
signicant
p
<
0
.
0001
and
were
incorporated
into
the
predicti
v
e
model.
The
performance
of
the
XGBoost
model
w
as
reected
in
an
area
under
the
recei
v
er
operating
characteristic
curv
e
(A
UR
OC)
of
0.87,
with
both
sensiti
vity
and
specicity
equal
to
0.77,
demonstrating
a
rob
ust
capacit
y
to
identify
risk
f
actors
related
Int
J
Artif
Intell,
V
ol.
15,
No.
4,
August
2026:
3888–3902
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
3891
to
sleep
disorders
[21],
[22].
In
a
dif
ferent
study
,
Haque
et
al.
[23]
created
a
predicti
v
e
model
to
assess
the
connection
between
academic
achie
v
ement
and
sleeplessness
in
kids
and
teenagers
bet
ween
the
ages
of
4
and
17.
F
or
this
purpose,
the
“Y
oung
minds
matter”
database
w
as
used,
applying
the
Boruta
algorithm
in
combination
with
the
RF
classier
to
select
the
most
rele
v
ant
features.
Subsequently
,
the
tree-based
pipeline
optimization
tool-classier
(TPO
TClassier)
w
as
used
to
identify
the
best-performing
supervised
models,
e
v
aluating
the
RF
,
XGBoost,
decision
tree,
and
Gaussian
nai
v
e
Bayes
algorithms.
As
a
result,
the
RF
model
demonstrated
the
best
performance,
achie
ving
a
precision
of
99%
and
an
accurac
y
of
95%,
outperforming
the
other
e
v
aluated
models
and
sho
wing
greater
ef
fecti
v
eness
in
predicting
insomnia
in
relation
to
academic
performance.
2.2.
Theor
etical
bases
2.2.1.
Extr
eme
gradient
boosting
XGBoost
is
dened
as
a
supervised
ML
algorithm
based
on
decision
trees.
Its
main
goal
is
to
perform
classication
and
re
gression
tasks
with
a
high
le
v
el
of
accurac
y
.
It
is
based
on
the
idea
of
gradient
boosting,
a
method
that
entails
training
models
one
after
the
other
,
with
each
ne
w
model
xing
the
mistak
es
of
the
preceding
ones.
What
mak
es
XGBoost
particularly
ef
fecti
v
e
is
that
it
is
optimized
to
deli
v
er
high
speed,
performance,
and
computational
ef
cienc
y
.
One
of
its
most
notable
features
is
its
ability
to
handle
lar
ge
v
olumes
of
data
quickly
,
thanks
to
a
highly
optimized
internal
architecture
that
includes
parallel
processing,
ef
cient
memory
usage,
and
adv
anced
tree
pruning
techniques
[24],
[25].
In
addition,
XGBoost
incorporates
re
gularization
mechanisms
(L1
and
L2
penalties
),
which
help
pre
v
ent
o
v
ertting
and
enhance
the
model’
s
generalization
ability
when
applied
to
ne
w
data.
Another
signicant
adv
antage
is
its
abili
ty
to
automatically
handle
missing
v
alues,
making
it
a
rob
ust
and
practical
tool
for
real-w
orld
applications
where
incomplete
data
is
common.
The
algorithm
also
allo
ws
ne-tuning
of
multiple
h
yperparameters,
such
as
tree
depth,
learning
rate,
and
the
number
of
estimators,
pro
viding
a
high
de
gree
of
customization
and
f
acilitating
model
optimization
according
to
the
specic
problem
[26],
[27].
2.2.2.
Sleep
quality
Sleep
quality
is
a
multidimensional
construct
that
refers
to
the
ef
cienc
y
,
continuity
,
depth,
and
subjecti
v
e
perception
of
nighttime
rest.
Medically
,
it
is
assessed
by
considering
multiple
objecti
v
e
and
subjecti
v
e
parameters,
including
sleep
latenc
y
,
sleep
ef
cienc
y
,
nocturnal
a
w
ak
enings,
total
sleep
duration,
and
the
presence
or
absence
of
sleep
disorders
such
as
insomnia,
sleep
apnea,
or
periodic
limb
mo
v
ements.
From
a
clinical
standpoint,
health
y
sleep
is
typically
dened
as
taking
no
longer
than
about
30
minutes
to
f
all
asleep
and
maintaining
a
sleep
ef
cienc
y
the
percentage
of
tim
e
actually
sleeping
while
in
bed
abo
v
e
85%,
a
total
duration
of
7
to
9
hours
per
night
in
adults,
and
a
r
educed
number
of
prolonged
nocturnal
a
w
ak
enings.
Additionally
,
the
indi
vidual’
s
subjecti
v
e
perception
upon
w
aking-such
as
feeling
rested
and
alert
is
a
rele
v
ant
component
in
the
clinical
assessment
of
sleep
quality
[28],
[29].
Sleep
quality
is
closely
related
to
multiple
neuroph
ysiological
functions,
including
memory
consolidation,
emotional
re
gulation,
cellular
repair
,
and
endocrine
balance.
Disruptions
in
sleep
ha
v
e
been
link
ed
to
an
increased
lik
elihood
of
psychiatric
conditions
including
major
depressi
v
e
disorder
and
anxiety
as
well
as
long-term
medical
problems
such
as
h
ypertension,
type
2
diabetes,
and
cardio
v
ascular
disease.
F
or
this
reason,
e
v
aluating
sleep
quality
is
considered
a
k
e
y
element
both
in
clinical
settings
and
in
research
on
general
health
[30].
Blue
light,
with
w
a
v
elengths
between
450
and
495
nanometers,
is
emitted
both
by
the
sun
and
by
electronic
de
vi
ces
such
as
cell
phones,
computer
s,
and
tele
visions.
While
benecial
during
the
day
,
i
ts
e
xposure
at
night
can
ne
g
ati
v
ely
af
fect
sleep.
The
circadian
rh
ythm,
which
re
gulates
the
sleep-w
ak
e
c
ycle,
is
controlled
by
the
suprachiasmatic
nucleus
(SCN)
of
the
h
ypothalamus
and
responds
primarily
to
ambient
light.
Melatonin,
a
k
e
y
hormone
for
inducing
sleep,
is
secreted
in
darkness,
and
its
production
decreases
when
e
xposed
to
light.
When
the
retina
detects
blue
light
at
night,
specialized
cells
(intrinsically
photosensiti
v
e
retinal
g
anglion
cells
(ipRGCs)
containing
melanopsin)
are
acti
v
ated
and
send
signals
to
the
SCN
indicating
that
light
is
still
present.
This
inhibits
melatonin
secretion,
delays
sleep
onset,
and
disrupts
the
circadian
rh
ythm.
As
a
result,
using
blue
light-emitting
de
vices
before
bedtime
can
lead
to
insomnia,
f
atigue,
and
a
decrease
in
sleep
quality
[31],
[32].
It
is
important
to
raise
a
w
areness
of
this
issue,
as
the
majority
of
people
use
electronic
de
vices
e
xcessi
v
ely
,
often
ignoring
the
health
risks
associated
with
blue
light
e
xposure.
Pr
edictive
model
based
on
mac
hine
learning
to
identify
sleep-r
elated
...
(Laberiano
Andr
ade-Ar
enas)
Evaluation Warning : The document was created with Spire.PDF for Python.
3892
❒
ISSN:
2252-8938
3.
METHOD
3.1.
Denition
of
the
kno
wledge
disco
v
ery
in
databases
methodology
The
KDD
methodology
is
an
iterati
v
e
and
interacti
v
e
process
that
inte
grates
e
xpert
kno
wledge
about
a
specic
problem
with
a
di
v
erse
set
of
data
analysis
techniques,
particularly
those
based
on
ML
algorithms.
This
process
consists
of
v
e
fundamental
stages:
selection,
preprocessing,
transformation,
data
mining,
and
interpretation/e
v
aluation
of
results.
Figure
1
sho
ws
that
each
of
these
phases
is
articulated
in
a
sequential
and
logical
manner
,
allo
wing
input
data
to
be
cleaned,
transformed,
and
subsequently
analyzed
in
order
to
e
xtract
useful
patterns
or
kno
wledge.
Finally
,
the
results
are
interpreted
and
e
v
aluated
according
to
the
initial
objecti
v
es
[33].
One
of
the
main
adv
antages
of
this
approach
is
its
e
xibility
,
as
it
allo
ws
returning
to
pre
vious
stages
of
the
process
if
the
results
are
unsatisf
actory
or
if
it
becomes
necessary
to
adjust
certain
assumptions.
This
ensures
greater
accurac
y
and
coherence
in
the
e
xtraction
of
kno
wledge
from
lar
ge
v
olumes
of
data.
Ultimately
,
this
methodology
represents
a
rob
ust
and
adaptable
frame
w
ork
for
v
arious
projects
and
promotes
continuous
impro
v
ement
thanks
to
its
iterati
v
e
nature.
Figure
1.
KDD
methodology
3.1.1.
Selection
In
this
section,
an
e
xhausti
v
e
search
w
as
conducted
on
the
Kaggle
platform,
a
widely
used
repository
within
the
scientic
and
technological
community
due
to
its
v
ast
v
ariety
of
public
datasets
from
dif
ferent
elds
of
kno
wledge.
During
the
search
process,
specic
criteria
were
emplo
yed
to
ensure
the
rele
v
ance
and
quality
of
the
selected
dataset.
K
e
yw
ords
such
as
sleep
health,
sleep
disorders,
sleep
quality
,
mental
health,
among
others,
were
used
within
Kaggle’
s
search
engine
to
nd
databases
containing
rele
v
ant
v
ariables
related
to
sleep
habits,
associated
disorders,
and
their
possible
connection
to
other
dimensions
of
ph
ysical
and
mental
health
[34],
[35].
Lik
e
wise,
priority
w
as
gi
v
en
t
o
selecting
a
dataset
in
CSV
format
due
to
its
tab
ular
structure,
compatibility
with
statistical
analysis
tools
(such
as
Python),
and
ease
of
manipulation.
This
format
allo
ws
for
ef
cient
data
processing,
f
acilitating
cleaning,
se
gmentation,
and
descripti
v
e
or
inferential
analysis
according
to
the
objecti
v
es
of
the
study
.
3.1.2.
Pr
epr
ocessing
In
the
preprocessing
stage,
descripti
v
e
statistics
re
v
eal
that
the
374
participants
ha
v
e
a
mean
age
of
42.18
years
(SD
=
8.67;
range
=
27–59),
with
an
a
v
era
g
e
sleep
duration
of
7.13
hours
(SD
=
0.80;
Q1
=
6.40;
median
=
7.20;
Q3
=
7.80;
minimum
=
5.80;
maximum
=
8.50),
and
an
a
v
erage
sleep
quality
score
of
7.31
(SD
=
1.20;
range
=
4–9).
The
a
v
erage
ph
ysical
acti
vity
le
v
el
is
59.17
(SD
=
20.83;
Q1
=
45;
median
=
60;
Q3
=
75;
range
=
30–90),
and
the
mean
stress
le
v
el
is
5.39
(SD
=
1.77;
range
=
3–8).
The
a
v
erage
heart
rate
is
70.17
bpm
(SD
=
4.14;
Q1
=
68;
median
=
70;
Q3
=
72;
range
=
65–86),
and
the
a
v
erage
daily
steps
are
6,816.84
(SD
=
1,617.92;
Q1
=
5,600;
median
=
7,000;
Q3
=
8,000;
range
=
3,000–10,000),
as
sho
wn
in
T
able
1.
Finally
,
all
numerical
v
ariables
were
standardized
using
the
StandardScaler()
function,
and
this
information
is
presented
in
T
able
2.
Int
J
Artif
Intell,
V
ol.
15,
No.
4,
August
2026:
3888–3902
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
3893
T
able
1.
Descripti
v
e
statistics
of
the
dataset
Statistics
Age
Sleep
duration
Quality
of
sleep
Ph
ysical
acti
vity
le
v
el
Stress
le
v
el
Heart
rate
Daily
steps
Count
374.00
374.00
374.00
374.00
374.00
374.00
374.00
Mean
42.18
7.13
7.31
59.17
5.39
70.17
6816.84
SD
8.67
0.80
1.20
20.83
1.77
4.14
1617.92
Min
27.00
5.80
4.00
30.00
3.00
65.00
3000.00
Q1
(25%)
35.25
6.40
6.00
45.00
4.00
68.00
5600.00
Q2
(Median)
43.00
7.20
7.00
60.00
5.00
70.00
7000.00
Q3
(75%)
50.00
7.80
8.00
75.00
7.00
72.00
8000.00
Max
59.00
8.50
9.00
90.00
8.00
86.00
10000.00
T
able
2.
T
echnical
steps
for
data
preprocessing
P
assed
T
echnical
operation
1
Con
v
ersion
to
num
eric:
pd.to
numeric(df[’Age’],
errors=’coerce’)
2
Imputat
ion
of
N
A:
df[’Sleep
Duration’].fillna(median,
inplace=True)
3
Calcul
ate
IQR
para
Heart
Rate
:
IQR
=
Q
3
−
Q
1
Outliers
=
{
x
<
Q
1
−
1
.
5
×
IQR
}
∪
{
x
>
Q
3
+
1
.
5
×
IQR
}
4
Standardi
zation:
StandardScaler().fit
transform(df[column
numbers])
3.1.3.
T
ransf
ormation
During
the
data
transformation
phase,
se
v
eral
preprocessing
techniques
were
applied
to
prepare
the
health
and
lifestyle
dataset
for
analysis.
Data
con
v
ersion
and
standardization
were
performed
e
xtensi
v
ely
,
ensuring
that
the
v
ariables
were
transformed
into
a
consistent
format
suitable
for
model
training.
Additional
procedures,
including
missing
v
alue
imputation
and
duplicate
record
remo
v
al,
were
required
only
to
a
limited
e
xtent,
while
outlier
detection,
v
ariable
transformation,
and
feature
engineering
were
applied
selecti
v
ely
to
impro
v
e
data
quality
.
Data
visualization
also
played
an
important
role
in
e
xploring
the
dataset,
together
with
the
analysis
of
the
v
ariance
e
xplained
by
the
principal
components,
as
illustrated
in
Figure
2.
Specically
,
Figure
2(a)
summarizes
the
preprocessing
w
orko
w
,
sho
wing
that
standardization
(100%)
and
data
con
v
ersion
(90%)
were
the
most
e
xtensi
v
ely
applied
techniques,
follo
wed
by
data
visualization
(80%)
and
outlier
detection
using
the
interquartile
range
(IQR)
method
(70%).
In
contrast,
missing
v
alue
imputation
(7.7%)
w
as
only
occasionally
necessary
,
and
no
duplicate
records
(0%)
were
detected,
highlighting
the
high
quality
of
the
dataset
prior
to
preprocessing.
Figure
2(b)
illustrates
the
tw
o-dimensional
representation
generated
using
principal
component
analysis
(PCA).
The
rst
tw
o
principal
components
e
xplain
93.4%
of
the
total
v
ariance
(PC1
=
65.2%
and
PC2
=
28.2%),
re
v
ealing
a
clear
separation
between
the
sleep
apnea
and
insomnia
classes
with
minimal
o
v
erlap.
This
distrib
ution
suggests
that
the
e
xtracted
features
ef
fecti
v
ely
preserv
e
the
underlying
class
structure,
pro
viding
a
solid
foundation
for
the
subsequent
classication
task.
(a)
(b)
Figure
2.
Data
manipulation
and
transformation
of
(a)
percentages
of
features
in
the
data
and
(b)
percentage
of
v
ariance
in
components
Pr
edictive
model
based
on
mac
hine
learning
to
identify
sleep-r
elated
...
(Laberiano
Andr
ade-Ar
enas)
Evaluation Warning : The document was created with Spire.PDF for Python.
3894
❒
ISSN:
2252-8938
3.1.4.
Data
mining
During
the
data
mining
procedure,
the
required
XGBoost
algorithm
w
as
applied.
The
results
in
Figure
3
indicate
that
the
XGBoost
model
demonstrates
e
xceptional
performance
according
to
the
precision-recall
and
recei
v
er
operating
characteristic
(R
OC)
curv
es.
The
precision-recall
curv
e,
with
an
a
v
erage
precision
of
0.95,
indicates
that
the
model
maintains
high
precision
e
v
en
with
a
high
recall,
making
it
highly
ef
fecti
v
e
at
identifying
positi
v
e
classes,
as
in
Figure
3(a).
Simultaneously
,
the
R
OC
curv
e,
with
an
A
UC
of
0.94,
conrms
e
xcellent
discriminati
v
e
ability
,
meaning
that
the
model
is
v
ery
capable
of
distinguishing
between
positi
v
e
and
ne
g
ati
v
e
classes
across
a
wide
range
of
thresholds,
as
in
Figure
3(b).
T
ogether
,
these
results
suggest
a
rob
ust
and
reliable
classier
for
the
task
at
hand.
Both
plots
demonstrate
that
the
XGBoost
model
performs
e
xceptionally
well
for
binary
class
ication
tasks.
The
precision-recall
curv
e,
with
an
a
v
erage
pre
cision
of
0.95,
sho
ws
that
the
model
maintains
high
precision
(fe
w
incorrect
positi
v
e
predictions)
e
v
en
while
achie
ving
high
recall
(detecting
most
actual
positi
v
e
cases).
Complementarily
,
the
R
OC
curv
e,
with
an
A
UC
of
0.94,
conrms
the
model’
s
e
xcellent
capacity
to
distinguish
between
positi
v
e
and
ne
g
ati
v
e
classes
o
v
er
a
broad
range
of
thresholds,
signicantly
outperforming
a
random
classier
[36],
[37].
(a)
(b)
Figure
3.
Classier
performance
metrics
for
XGBoost
model:
(a)
precision-recall
curv
e
and
(b)
R
OC
curv
e
4.
MA
THEMA
TICAL
B
ASIS
OF
THE
EXTREME
GRADIENT
BOOSTING
MODEL
The
follo
wing
section
presents
the
mathematical
foundation
of
the
XGBoost
algorithm,
adapted
to
the
binary
classication
problem
aimed
at
identifying
health-related
conditions
associated
with
sle
ep.
It
describes
the
k
e
y
concepts
and
equations
that
go
v
ern
the
model’
s
lea
rning
process.
The
section
also
e
xplains
ho
w
the
algorithm
optimizes
its
objecti
v
e
function
to
impro
v
e
predicti
v
e
performance.
4.1.
General
objecti
v
e
function
XGBoost
b
uilds
an
ensemble
of
additi
v
e
decision
trees
by
optimizing
an
objecti
v
e
function
that
combines
a
loss
function
with
a
re
gularization
term
to
pre
v
ent
o
v
ertting
as
(1)
[38].
This
approach
allo
ws
the
model
to
s
equentially
correct
the
errors
of
pre
vious
trees.
It
also
balances
model
comple
xity
and
predicti
v
e
accurac
y
to
achie
v
e
rob
ust
performance
on
unseen
data.
L
(
t
)
=
n
X
i
=1
l
(
y
i
,
ˆ
y
(
t
)
i
)
+
t
X
k
=1
Ω(
f
k
)
(1)
Where
l
(
y
i
,
ˆ
y
(
t
)
i
)
is
loss
function
between
the
prediction
and
the
actual
label;
and
Ω(
f
k
)
is
penalty
on
tree
comple
xity
f
k
.
Int
J
Artif
Intell,
V
ol.
15,
No.
4,
August
2026:
3888–3902
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
3895
4.2.
Model
r
egularization
Re
gularization
controls
the
comple
xity
of
the
model
by
penalizi
ng
the
depth
and
the
weight
of
the
lea
v
es
in
each
tree
as
in
(2)
[39].
This
helps
pre
v
ent
o
v
ertting
by
discouraging
o
v
erly
comple
x
trees.
It
ensures
that
the
model
generalizes
well
to
ne
w
,
unseen
data.
Ω(
f
)
=
γ
T
+
1
2
λ
T
X
j
=1
w
2
j
(2)
Where
T
:
number
of
lea
v
es;
w
j
is
predicted
v
alue
on
the
sheet
j
;
γ
is
penalty
for
each
leaf;
and
λ
is
L2
re
gularization
parameter
.
4.3.
Logistic
loss
function
F
or
binary
classication,
the
l
ogistic
loss
(log-loss)
is
used
[40],
which
measures
the
error
between
the
predicted
probability
and
the
actual
class,
as
in
(3).
This
loss
function
penalizes
incorrect
predictions
more
hea
vily
when
the
model
is
condent
b
ut
wrong.
It
pro
vides
a
smooth
gradient
that
f
acilitates
optimization
during
the
training
of
XGBoost.
l
(
y
i
,
ˆ
y
i
)
=
−
[
y
i
log
(
σ
(
ˆ
y
i
))
+
(1
−
y
i
)
log
(1
−
σ
(
ˆ
y
i
))]
(3)
Where
the
sigmoid
function
is:
σ
(
ˆ
y
i
)
=
1
1
+
e
−
ˆ
y
i
4.4.
Optimization
by
second
order
expansion
T
o
impro
v
e
training
ef
cienc
y
,
XGBoost
uses
a
second-order
(T
aylor)
e
xpansion
[41]
around
the
current
prediction,
as
in
(4).
This
approximation
allo
ws
the
model
to
consider
both
the
gradient
and
the
Hessian
when
updating
tree
lea
v
es.
It
accelerat
es
con
v
er
gence
and
enhances
the
accurac
y
of
each
iteration
during
training.
L
(
t
)
≈
n
X
i
=1
g
i
f
t
(
x
i
)
+
1
2
h
i
f
t
(
x
i
)
2
+
Ω(
f
t
)
(4)
Where
g
i
=
∂
l
(
y
i
,
ˆ
y
(
t
−
1)
i
)
∂
ˆ
y
(
t
−
1)
i
is
gradient
and
h
i
=
∂
2
l
(
y
i
,
ˆ
y
(
t
−
1)
i
)
∂
ˆ
y
(
t
−
1)2
i
is
Hessian.
4.5.
Optimal
weight
of
a
sheet
The
optimal
v
alue
that
a
leaf
should
predict
is
calculated
as
(5).
This
v
alue
is
deri
v
ed
by
minimizing
the
re
gularized
objecti
v
e
function
for
that
specic
leaf.
It
ensures
that
each
leaf
contrib
utes
optimally
to
reducing
the
o
v
erall
prediction
error
of
the
model.
w
∗
j
=
−
G
j
H
j
+
λ
(5)
Where
G
j
=
P
i
∈
I
j
g
i
is
sum
of
gradients;
H
j
=
P
i
∈
I
j
h
i
is
sum
of
Hessians;
and
I
j
is
set
of
samples
o
n
the
sheet
j
.
4.6.
Pr
ot
fr
om
a
di
vision
T
o
decide
whether
to
perform
a
split
at
a
node,
the
resulting
information
g
ain
is
calculated
(6).
Thi
s
g
ain
measures
ho
w
much
the
split
reduces
the
o
v
erall
loss
of
the
model.
It
helps
the
algorithm
choose
the
most
informati
v
e
splits,
impro
ving
predicti
v
e
accurac
y
while
controlling
model
comple
xity
.
Gain
=
1
2
G
2
L
H
L
+
λ
+
G
2
R
H
R
+
λ
−
(
G
L
+
G
R
)
2
H
L
+
H
R
+
λ
−
γ
(6)
Where
G
L
,
H
L
is
sum
of
gradients
and
Hessians
to
the
left
of
the
split;
G
R
,
H
R
is
sum
to
the
right
of
the
split;
and
γ
is
penalty
for
creating
an
additional
sheet.
Pr
edictive
model
based
on
mac
hine
learning
to
identify
sleep-r
elated
...
(Laberiano
Andr
ade-Ar
enas)
Evaluation Warning : The document was created with Spire.PDF for Python.
3896
❒
ISSN:
2252-8938
5.
RESUL
TS
5.1.
Ev
aluation
of
r
esult
Figure
4
compares
three
classication
models
(logistic
re
gression,
RF
,
and
XGBoost),
sho
wing
that
all
three
e
xhibit
e
xcellent
and
comparable
performance
[42].
The
R
OC
curv
e
demonstrates
that
all
models
are
equally
capable
of
distinguishing
between
classes
(A
UC
=
0.95),
as
sho
wn
in
Figure
4(a).
The
precision-recall
curv
e
indicates
that
XGBoost
is
slightly
superior
in
maintaining
a
balance
between
precision
and
recall
(AP
=
0.92),
as
sho
wn
in
Figure
4(b),
while
the
cumulati
v
e
g
ain
chart
conrms
that
all
models
are
highly
ef
cient
at
quickly
identifying
the
majority
of
positi
v
e
cases,
as
sho
wn
in
Figure
4(c).
(a)
(b)
(c)
Figure
4.
Model
performance
comparison:
(a)
logistic
re
gression,
(b)
RF
,
and
(c)
XGBoost
Furthermore,
other
results
sho
wn
in
Figure
5
include
normalized
confusion
matrices
for
the
three
ML
models
e
v
aluated
for
the
binary
classication
task,
allo
wing
for
a
detailed
comparison
between
the
predicted
and
actual
class
labels.
Figure
5(a)
sho
ws
the
performance
of
the
logistic
re
gression
model,
which
correctly
classied
91%
of
the
class
0
samples
and
94%
of
the
class
1
samples,
with
misclassication
rates
of
9%
and
6%,
respecti
v
ely
.
On
the
other
hand,
Figure
5(b)
illustrates
the
results
obtained
with
the
RF
classier
,
which
achie
v
ed
corr
ect
classication
rates
of
88%
for
class
0
and
94%
for
class
1,
although
it
had
a
slightly
higher
f
alse
positi
v
e
rate
(12%)
than
the
other
models.
Finally
,
Figure
5(c)
presents
the
confusion
matrix
for
the
XGBoost
model,
whi
ch
of
fered
the
best
o
v
erall
performance
by
correctly
identifying
97%
of
class
0
instances
Int
J
Artif
Intell,
V
ol.
15,
No.
4,
August
2026:
3888–3902
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
3897
and
94%
of
class
1
instances,
while
reducing
the
f
alse
positi
v
e
rate
to
3%
and
maintaining
a
lo
w
f
alse
ne
g
ati
v
e
rate
(6%).
Ov
erall,
the
results
indicate
that
all
three
classiers
achie
v
ed
high
predicti
v
e
performance,
with
XGBoost
demonstrating
the
most
accurate
and
balanced
classication.
(a)
(b)
(c)
Figure
5.
Confusion
matrices
of
classication
models:
(a)
logistic
re
gression,
(b)
RF
,
and
(c)
XGBoost
T
able
3
reports
t
he
o
v
erall
error
metrics.
The
model
sho
ws
a
lo
w
f
alse
positi
v
e
rate
(FPR)
of
0.065
and
a
true
positi
v
e
rate
(TPR)
of
0.909,
indicating
high
o
v
erall
sensiti
vity
.
The
true
ne
g
ati
v
e
rate
(TNR)
is
also
high
(0.935),
reecting
good
specicity
.
The
o
v
erall
error
rate
is
relati
v
ely
lo
w
(0.085),
which
supports
the
model’
s
general
reliability
.
T
able
3.
Global
error
metrics
Metric
V
alue
FPR
0.065
FNR
0.182
TPR
0.909
TNR
0.935
Error
rate
0.085
The
class-wise
metrics
presented
in
T
able
4
reect
a
balanced
performance
of
the
model
in
classi
fying
insomnia
and
sleep
apnea.
The
insomnia
class
sho
ws
a
precision
of
0.862,
a
perfect
recall
of
1.000,
and
an
F1-score
of
0.926,
indicating
that
all
i
nsomnia
cases
were
correctly
identied
with
no
f
alse
ne
g
ati
v
es.
Pr
edictive
model
based
on
mac
hine
learning
to
identify
sleep-r
elated
...
(Laberiano
Andr
ade-Ar
enas)
Evaluation Warning : The document was created with Spire.PDF for Python.