IAES
Inter
national
J
our
nal
of
Articial
Intelligence
(IJ-AI)
V
ol.
15,
No.
3,
June
2026,
pp.
2797
∼
2810
ISSN:
2252-8938,
DOI:
10.11591/ijai.v15.i3.pp2797-2810
❒
2797
Deep
h
ybrid
models
f
or
bitcoin
f
or
ecasting:
EMD,
CEEMD
AN,
and
LSTM
in
comparison
A
y
oub
Aarabi,
Mary
em
Ait
Moulay
,
Issam
Bouganssa,
Abdelali
Lasfar
Laboratory
Systems
Analysis,
Information
Processing,
and
Industrial
Management,
High
School
of
T
echnology
Sal
´
e,
Mohammed
V
Uni
v
ersity
in
Rabat,
Rabat,
Morocco
Article
Inf
o
Article
history:
Recei
v
ed
Jul
28,
2025
Re
vised
Apr
28,
2026
Accepted
May
11,
2026
K
eyw
ords:
Bitcoin
Complete
ensemble
empirical
mode
decomposition
with
adapti
v
e
noise
Deep
learning
Long
short-term
memory
Machine
learning
ABSTRA
CT
In
this
study
,
an
articial
neural
netw
ork
(ANN)
w
as
de
v
eloped
to
forecast
Bitcoin
prices
using
one
of
the
most
successful
deep
learning
architectures
for
time
series
analysis:
long
short-term
memory
(LSTM)
netw
orks.
This
model
w
as
enha
nced
with
a
signal
processing
layer
that
reduces
the
impact
of
the
instrument’
s
high
v
olatility
on
prediction
accur
ac
y
by
applying
tw
o
signal
decomposition
techniques:
empiri
cal
mode
decomposition
(EMD)
and
complete
ensemble
empirical
mode
decomposition
with
adapti
v
e
noise
(CEEMD
AN).
This
study
is
moti
v
ated
by
the
major
uctuations
in
Bitcoin
prices,
which
mak
e
precise
forecasting
dif
cult
b
ut
crucial
for
e
xperts
and
in
v
estors.
This
ndings
demonstrate
that
forecasting
performance
impro
v
es
when
decomposition
techniques
are
used.
In
particular
,
compared
to
the
con
v
entional
LSTM
and
EMD-LSTM
models,
the
CEEMD
AN-LSTM
model
achie
v
ed
the
highest
a
ccurac
y
,
with
a
mean
absolute
error
(MAE)
of
167.837
and
a
root
mean
square
error
(RMSE)
of
255.673,
outperforming
both
EMD-LSTM
(MAE
=168.785,
RMSE
=256.042)
and
the
standard
LSTM
(MAE
=169.516,
RMSE
=256.225).
The
combination
of
CEEMD
AN
and
LSTM
results
i
n
a
more
reliable
model
that
can
accurately
capture
short-term
uctuations
i
n
Bitcoin
prices.
This
is
an
open
access
article
under
the
CC
BY
-SA
license
.
Corresponding
A
uthor:
A
youb
Aarabi
Laboratory
Systems
Analysis,
Information
Processing,
and
Industrial
Management
High
School
of
T
echnology
Sal
´
e,
Mohammed
V
Uni
v
ersity
in
Rabat
Rabat,
Morocco
Email:
ayoub
.arb@gmail.com
1.
INTR
ODUCTION
The
Ne
w
Y
ork
Stock
Exchange
(NYSE),
the
w
orld’
s
lar
gest
stock
e
xchange,
is
currently
c
o
ns
idering
a
transition
to
w
ard
e
xtended
or
e
v
en
continuous
trading
hours.
This
shift
reects
broader
changes
in
global
nancial
mark
ets,
where
cryptocurrencies,
led
by
Bitcoin,
ha
v
e
introduced
a
model
of
uninterrupted
trading.
In
this
conte
xt,
traditional
mark
et
structures
may
appear
increasingly
misaligned
with
continuously
operating
digital-asset
en
vironments.
Bitcoin
w
as
introduced
by
Nakamoto
in
a
seminal
white
paper
published
in
2008
[1],
with
the
objecti
v
e
of
enabling
a
decentralized
digital
payment
system
without
reliance
on
intermediaries.
It
is
bas
ed
on
blockchain
technology
,
which
ensures
transparenc
y
and
resistance
to
tampering
through
a
proof-of-w
ork
consensus
mechanism
[2],
[3].
Each
participant
in
the
netw
ork
maintains
a
complete
cop
y
of
the
transaction
ledger
,
ensuring
rob
ustness,
v
eriability
,
and
transparenc
y
[4].
Initially
re
g
arded
with
sk
epticism,
Bitcoin
has
e
xperienced
signicant
gro
wth,
with
its
ma
rk
et
J
ournal
homepage:
http://ijai.iaescor
e
.com
Evaluation Warning : The document was created with Spire.PDF for Python.
2798
❒
ISSN:
2252-8938
capitalization
surpassing
$1
trillion
in
recent
years
and
increasing
institutional
adoption.
It
is
no
w
often
considered
a
potential
safe-ha
v
en
asset,
comparable
to
gold
[5],
and
a
hedge
ag
ainst
inationary
pressures
and
e
xpansi
onary
monetary
policies
[6].
Some
countries,
such
as
El
Salv
ador
and
the
Central
African
Republic,
ha
v
e
adopted
Bitcoin
as
le
g
al
tender
or
incorporated
it
i
nto
national
reserv
es
[7].
At
the
same
time,
major
central
banks,
including
the
European
Central
Bank
and
the
People’
s
Bank
of
China,
are
e
xploring
the
de
v
elopment
of
central
bank
digital
currencies
(CBDCs)
[8],
[9].
This
gro
wing
recognition
highlights
the
strate
gic
importance
of
cryptocurrencies
and
moti
v
ates
further
in
v
estig
ation
into
their
v
olatile
dynamics.
Unlik
e
traditional
nancial
assets,
cryptocurrenc
y
prices
are
inuenced
by
a
combination
of
technological,
macroeconomic,
and
beha
vioral
f
actors,
which
mak
es
their
modeling
particular
ly
challenging
[10].
In
this
conte
xt,
de
v
eloping
reliable
forecasting
models,
especially
those
based
on
deep
learning,
has
become
an
important
scientic
and
economic
issue.
2.
RELA
TED
W
ORKS
Price
prediction
of
cryptocurrencies
is
an
inherently
dif
cult
task,
because
it
in
v
olv
es
fore
casting
future
v
alues
of
a
long
and
highly
v
olatile
nancial
time
series.
This
com
p
l
e
xity
is
the
result
of
v
ery
high
v
olatility
of
cryptocurrenc
y-based
nancial
time
series,
which
are
inuenced
by
economic
,
political,
technological,
and
psychological
forces.
Such
ser
ies
are
typically
non-stationary
and
nonlinear
,
and
are
frequently
e
xposed
to
e
xogenous
shocks.
T
o
address
the
limitations
of
the
classical
methods
that
f
ail
to
capture
the
dynamics
of
these
series,
man
y
recent
studies
rely
on
deep
learning
approaches.
In
this
re
g
ard,
Derbentse
v
et
al.
[11]
in
v
estig
ated
the
short-term
dynamics
of
Bitcoin,
Ethereum,
and
Ripple
using
sophisticated
algorithms
including
articial
neural
netw
orks
(ANNs),
random
forests
(RFs)
and
binary
autore
gressi
v
e
trees
(B
AR
Ts).
The
y
found
that
both
ANN
and
B
AR
T
models
substantially
outperformed
nai
v
e
classication
methods
(up
to
63%
accurac
y
out-of-sample)
based
on
more
than
1,500
daily
observ
ations
between
2015
and
2019.
Cho
wdhury
et
al.
[12]
continued
with
a
lar
ger
number
of
cryptocurrencies,
including
the
CCI30
inde
x,
and
used
more
adv
anced
machine
learning
models
as
boosted
trees,
k-nearest
neighbors,
and
rob
ust
ensemble
models.
The
authors’
research
for
the
years
2017
to
2019
sho
wed
that
their
ensemble
models
and
their
boosted
tree
methods
performed
v
ery
well,
surpassing
e
v
en
se
v
eral
sophisticated
models.
Ho
we
v
er
,
multiple
studies
ha
v
e
since
underscored
the
inherent
limitations
of
deep
learning
models
when
the
y
are
applied
to
this
kind
of
time
series.
Pintelas
et
al.
[13]
sho
wed
that
long
short-term
memory
(LSTM)-
or
con
v
olutional
neural
netw
ork
(CNN)-based
models
can
be
po
werful,
their
performance
de
grades
in
the
presence
of
noise
and
missing
data.
F
or
a
more
rob
ust
prediction
and
a
prediction
with
a
higher
reliability
,
Li
vieris
et
al.
[14]
introduced
ensemble
architectures
using
bagging,
stacking,
and
a
v
eraging
methods
in
v
olving
sets
of
se
v
eral
LSTM
and
Con
v1D
learners
on
hourly
data
for
Bitcoin,
Ethereum,
and
Ripple.
The
obtained
results
suggest
that
the
h
ybrid
models,
though
computationally
e
xpensi
v
e,
e
xhibit
lar
ge
mar
gins
in
accurac
y
.
In
a
similar
,
more
specic
manner
,
P
atel
et
al.
[15]
proposed
a
h
ybrid
LSTM-g
ated
recurrent
unit
(GR
U)
model
for
Litecoin
and
Monero
price
prediction.
The
results
of
e
xperiments
demonstrate
that
this
recurrent
model
could
obtain
better
results
when
comparing
with
the
simple
LSTM
netw
orks,
especially
in
terms
of
root
mean
squared
error
(RMSE)
and
mean
absolute
percentage
error
(MAPE).
Other
studies,
such
as
in
[16],
[17],
emphasized
the
importance
of
e
xplanatory
v
ariable
selection,
emplo
ying
methods
such
as
the
granger
causality
,
mutual
information,
and
gre
y
relational
analysis
(GRA)
to
quantify
the
inue
n
c
e
of
macroeconomic
v
ariables
on
Bitcoin
price.
More
recent
w
orks
ha
v
e
compared
deep
h
ybrid
models
to
the
traditional
approaches.
Chauhan
et
al.
[18]
considered
a
CNN–GR
U
model
for
long
horizon
prediction
in
the
Bitcoin
price
and
found
that
deeper
netw
ork
structures
do
not
necessarily
lead
to
better
performance
without
proper
pre-processing
of
signals.
Similarly
,
Qureshi
et
al.
[19]
analyzed
some
classical
(e.g.,
auto-re
gressi
v
e
inte
grated
mo
ving
a
v
erage
(ARIMA),
multi-layer
perceptron
(MLP),
and
e
xtreme
learning
machine
(ELM))
and
h
ybrid
models,
and
the
y
noted
that
these
models
are
not
able
to
accurately
capture
the
nonlinear
dynamics
of
the
cryptocurrenc
y
price
e
v
olution,
especially
during
the
high
v
olatility
periods.
Ho
we
v
er
,
a
maj
or
limitation
of
all
these
studies
is
that
the
y
ha
v
e
generally
preferred
to
impro
v
e
the
predicti
v
e
performance
by
simply
e
xploiting
increasingly
comple
x
architectures,
b
ut
the
y
ha
v
e
paid
less
attention
to
the
solid
construction
of
training
sets
that
are
full
of
conte
xtual
information.
Nonetheless,
adding
appropriate
e
xogenous
v
ar
iables,
such
as
oil,
gold,
interest
rates,
or
dollar
strength
inde
x
es,
can
signicantly
impro
v
e
the
accurac
y
of
forecasts,
espec
ially
in
highly
v
olatile
cryptocurrenc
y
mark
ets.
This
g
ap
is
addressed
in
the
present
study
by
inte
grating
a
ne-grained
analysis
of
basic
economic
v
ariables,
where
GRA
is
emplo
yed
to
Int
J
Artif
Intell,
V
ol.
15,
No.
3,
June
2026:
2797–2810
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
2799
select
the
most
rele
v
ant
e
xplanatory
features
and
construct
an
economically
informati
v
e
and
coherent
input
set.
3.
METHOD
F
orecasting
Bitcoin
prices
is
a
highly
challenging
task
gi
v
en
the
high
v
olatility
,
non-st
ationary
,
and
nonlinear
nature
of
the
ass
et
[20],
[21].
T
o
impro
v
e
forecasting
results
under
such
conditions,
a
data
preprocessing,
decomposition,
and
sequential
learning
approach
w
as
proposed.
The
general
architecture
of
the
proposed
methodology
is
illustrated
in
Figure
1.
Under
this
methodology
v
arious
input
v
ariables
ha
v
e
been
tak
en
into
account
including
open,
high,
lo
w
,
close
prices,
v
olume
traded,
and
macroeconomic
indicators.
GRA
is
used
for
the
determination
of
the
most
inuenti
al
features
that
share
structural
similarities
with
Bitcoin
closing
price.
At
the
same
time,
the
tar
get
v
ariable
i.e.
the
closing
price
series,
w
ould
be
decomposed
using
EMD
or
its
impro
v
ed
v
ersion
kno
wn
as
complete
ensemble
empirical
mode
decomposition
with
adapti
v
e
noise
(CEEMD
AN).
This
process
yields
a
set
of
intrinsic
mode
functions
(IMF)
and
a
residual
term,
where
each
is
associated
with
a
particular
frequenc
y
component
of
the
signal.
Each
IMF
and
the
residue
are
modeled
separately
using
an
LSTM
netw
ork.
These
s
ub-models
capture
dif
ferent
temporal
characteristics
and
help
pre
v
ent
o
v
ert
ting
to
noise.
The
indi
vidual
forecasts
are
subsequently
recombined
through
a
reconstruction
process
to
gi
v
e
the
nal
forecast
of
closing
price
of
the
Bitcoin.
In
addition
to
short-term
v
olatility
reduction,
such
a
modular
architecture
also
enhances
the
interpretability
and
rob
ustness
of
the
predicti
v
e
system.
It
allo
ws
a
multi-resolution
representation
of
the
mark
et
dynamics
which
reects
the
inherent
heterogeneity
and
noisy
features
of
nancial
time
series.
Figure
1.
Architecture
of
our
inte
grated
forecasting
model
3.1.
Experimental
setup
and
r
epr
oducibility
The
dataset
(2014
to
2024,
daily)
is
split
chronologically
into
70%
training
(2018
to
2021),
15%
v
alidation
(2022),
and
15%
testing
(2023
to
2024)
to
a
v
oid
data
leakage.
One-step-ahead
forecasting
is
performed
using
a
sliding
input
windo
w
of
L
=
30
days.
All
features
are
normalized
using
Min–Max
scaling
tted
on
the
training
set.
T
raining
uses
the
Adam
optimizer
(learning
rate
10
−
3
)
and
Huber
loss.
The
LSTM
architecture
consists
of
tw
o
layers
(100
and
32
units)
with
dropout
0.2
and
early
stopping.
Models
are
trained
for
up
to
100
epochs
with
batch
size
32.
All
e
xperiments
are
im
plemented
in
Python
using
K
eras/T
ensorFlo
w
with
a
x
ed
random
seed.
Bot
h
EMD
and
CEEMD
AN
decompositions
yielded
v
e
IMFs
in
the
e
xperiments
(with
an
additional
residual
component
for
CEEMD
AN).
Each
component
w
as
modeled
by
an
independent
LSTM,
and
the
nal
predict
ion
w
as
obtained
by
aggre
g
ating
the
forecasts
of
all
components.
As
summarized
in
T
able
1,
the
main
h
yperparameters
are
reported
for
reproducibility
.
In
the
follo
wing,
the
three
forecasting
architectures
considered
in
this
study
are
detailed:
a
benchmark
LSTM
model,
an
EMD
based
h
ybrid
model,
and
a
CEEMD
AN-enhanced
v
ersion.
3.2.
Long
short-term
memory
on
nancial
time
series
The
LSTM
neural
netw
orks
[22]
are
commonly
used
for
time
series
predictions
tasks,
including
in
nancial
mark
ets
[23],
in
ener
gy
[24]
and
in
cryptocurrenci
es
[25].
The
y
surpass
classic
neural
netw
orks
Deep
hybrid
models
for
bitcoin
for
ecasting:
EMD,
CEEMD
AN,
and
LSTM
in
comparison
(A
youb
Aar
abi)
Evaluation Warning : The document was created with Spire.PDF for Python.
2800
❒
ISSN:
2252-8938
in
situations
where
memory
ef
fects
are
important
as
a
consequence
of
their
capacity
to
learn
long-term
dependencies
in
sequential
data.
The
memory
cell
C
t
in
the
LSTM
is
managed
using
input
i
t
,
for
get
f
t
,
and
output
o
t
g
ates
that
enable
the
information
in
the
ce
ll
to
be
k
ept
or
for
gotten
o
v
er
time,
which
mak
es
it
ideal
for
the
unpredictable
nature
of
the
cryptocurrenc
y
mark
et.
Ho
we
v
er
,
direct
application
of
LSTM
to
e
xtremely
noisy
data,
e.g.,
ra
w
Bitcoin
prices,
may
result
in
o
v
ertting
to
t
he
noise,
causing
unstable
prediction
[26].
At
each
time
step
t
,
gi
v
en
the
input
x
t
,
pre
vious
hidden
state
h
t
−
1
,
and
pre
vious
cell
state
C
t
−
1
,
the
LSTM
performs
computations
as
in
(1)
to
(6).
f
t
=
σ
(
W
f
·
[
h
t
−
1
,
x
t
]
+
b
f
)
(F
or
get
g
ate)
(1)
i
t
=
σ
(
W
i
·
[
h
t
−
1
,
x
t
]
+
b
i
)
(Input
g
ate)
(2)
˜
C
t
=
tanh(
W
C
·
[
h
t
−
1
,
x
t
]
+
b
C
)
(Candidate
cell
state)
(3)
C
t
=
f
t
∗
C
t
−
1
+
i
t
∗
˜
C
t
(Cell
state
update)
(4)
o
t
=
σ
(
W
o
·
[
h
t
−
1
,
x
t
]
+
b
o
)
(Output
g
ate)
(5)
h
t
=
o
t
∗
tanh(
C
t
)
(Hidden
state)
(6)
Here,
σ
denotes
the
sigmoid
acti
v
ation
function,
tanh
is
the
h
yperbolic
tangent
function,
and
∗
denotes
element-wise
multiplication.
These
mechanisms
allo
w
LSTMs
to
learn
comple
x
temporal
dynamics.
T
able
1.
Main
h
yperparameters
used
in
the
e
xperiments
Component
Setting
Look-back
windo
w
L
30
days
F
orecast
horizon
1-step
ahead
(
t
+
1
)
LSTM
layers
2
Units
per
layer
100,
32
Dropout
rate
0.2
Optimizer
Adam
Learning
rate
0.001
Batch
size
32
Epochs
100
(early
stopping
enabled)
Loss
function
Huber
loss
3.3.
EMD-LSTM:
fr
equency
di
vision
and
multiple
le
v
els
of
r
epr
esentation
T
o
o
v
ercome
this
v
olatility
,
we
utilize
a
h
ybrid
approach
EMD
and
LSTM.
Huang
et
al.
[27]
has
proposed
the
EMD,
an
y
time
series
is
decomposed
into
a
number
of
IMFs,
all
of
which
correspond
to
oscillations
ha
ving
dif
ferent
characteristi
c
time
scales.
Suc
h
a
decomposition
is
adapti
v
e
and
data-dri
v
en,
and
is
appropriate
for
nonlinear
and
nonstationary
signals,
and
mark
et
prices.
Each
obtained
IMF
,
c
i
(
t
)
,
is
presumed
to
ha
v
e
a
higher
le
v
el
of
re
gularity
compared
to
the
ra
w
series
and
in
turn
easier
for
modeling.
Therefore,
the
signal
can
be
written
as
in
(7).
x
(
t
)
=
n
X
i
=1
c
i
(
t
)
+
r
n
(
t
)
(7)
Where
r
n
(
t
)
is
the
non-oscillatory
residual.
Each
such
component
is
modeled
as
independent
LSTM
and
their
output
are
combined
to
generate
the
nal
prediction.
This
frequenc
y
decoupling
w
ould
mak
e
the
model
v
ery
rob
ust
in
terms
of
the
mark
et
state
[28].
T
o
capture
the
subtle
frequenc
y
patterns
embedded
in
the
original
Bitcoin
price
signal,
we
use
the
EMD,
a
data-dri
v
en
method
that
decomposes
a
non-stationary
and
non-linear
time
series
into
a
limited
set
of
IMFs.
The
decomposition
algorithm
is
gi
v
en
in
Algorithm
1.
The
Algorithm
1
decomposes
the
original
signal
adapti
v
ely
into
a
nite
number
of
IMFs,
each
representing
a
specic
local
oscillation
mode
of
the
original
signal.
In
the
case
of
Bitcoin
forecasting,
EMD
enables
the
short
term
uctuations
to
be
separated
from
long
term
trends,
making
it
possible
to
learn
more
stable
representations
of
both
using
separate
LSTM
models.
Ho
we
v
er
,
as
we
will
see
in
the
follo
wing
section,
EMD
is
vulnerable
to
mode
mixing,
which
dri
v
es
the
need
for
its
impro
v
ed
v
ersion.
Int
J
Artif
Intell,
V
ol.
15,
No.
3,
June
2026:
2797–2810
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
2801
Algorithm
1
EMD
1:
Input:
Signal
x
(
t
)
2:
Output:
Intrinsic
mode
functions
{
c
1
(
t
)
,
c
2
(
t
)
,
.
.
.
,
c
n
(
t
)
}
and
residual
r
n
(
t
)
3:
Initialize:
r
0
(
t
)
←
x
(
t
)
4:
f
or
i
=
1
to
n
do
5:
h
0
(
t
)
←
r
i
−
1
(
t
)
6:
while
h
k
(
t
)
is
not
an
IMF
do
7:
Identify
local
e
xtrema
of
h
k
(
t
)
8:
Interpolate
upper
and
lo
wer
en
v
elopes
9:
Compute
mean
en
v
elope
m
k
(
t
)
10:
h
k
+1
(
t
)
←
h
k
(
t
)
−
m
k
(
t
)
11:
end
while
12:
c
i
(
t
)
←
h
k
(
t
)
13:
r
i
(
t
)
←
r
i
−
1
(
t
)
−
c
i
(
t
)
14:
end
f
or
15:
r
etur
n
{
c
1
(
t
)
,
.
.
.
,
c
n
(
t
)
}
and
r
n
(
t
)
3.4.
CEEMD
AN-LSTM:
enhanced
r
ob
ustness
and
mode
mixing
a
v
oided
Ho
we
v
er
,
EMD
suf
fers
from
mode
mixing,
a
mixing
of
frequenc
y
scales
in
a
single
IMF
,
leading
to
obstacles
to
learning.
T
o
resolv
e
it,
the
CEEMD
AN
method
[29]
is
an
enhancement
performed
to
stabilize
frequenc
y
separat
ion.
CEEMD
AN
introduces
white
stochastic
Gaussian
Noise
to
n
copies
of
the
signal
in
a
noise-controlled
manner
,
apply
EMD
on
each
n,
and
a
v
erages
the
decompositions
to
get
smooth
and
stable
IMFs.
This
approach
results
in
impro
v
ed
component
orthogonality
and
true
reconstructability
while
remaining
resistant
to
structural
noise.
The
signal
can
be
written
as
in
(8).
x
(
t
)
=
n
X
i
=1
¯
c
i
(
t
)
+
r
n
(
t
)
(8)
Where
¯
c
i
(
t
)
is
the
a
v
erage
of
the
IMFs
e
xtracted
at
each
noisy
step.
CEEMD
AN
has
been
applied
in
v
arious
elds,
such
as
stock
inde
x
forecasting,
biomedical
signal
processing,
and
climate
analysis,
and
cryptocurrenc
y
price
modeling
[30].
When
combined
with
LSTM,
it
separates
out
the
high-frequenc
y
components
(noise
and
micro-oscillations)
to
enable
the
training
of
LSTM
to
be
concentrated
on
the
signals,
which
are
smoother
and
ha
v
e
more
information.
Therefore
the
CEEMD
AN-LSTM
netw
ork
i
s
especially
t
to
the
aforementioned
v
olatile
and
noisy
nancial
series.
T
o
o
v
ercome
EMD
weaknesses,
such
as
mode
mixing,
CEEMD
AN
is
utilized
t
o
pro
vide
rob
ust
decomposition.
The
algorithm
of
the
iterati
v
e
ensemble
approach
to
procure
noise-rob
ust
IMFs
from
the
input
signal
is
presented
in
Algorithm
2.
Algorithm
2
Complete
ensemble
EMD
with
adapti
v
e
noise
(CEEMD
AN)
1:
Input:
Signal
x
(
t
)
,
white
noise
w
(
t
)
,
ensemble
size
M
,
noise
le
v
el
ε
2:
Output:
Components
{
c
1
(
t
)
,
c
2
(
t
)
,
.
.
.
,
c
n
(
t
)
}
and
residual
r
n
(
t
)
3:
f
or
m
=
1
to
M
do
4:
x
m
(
t
)
←
x
(
t
)
+
ε
·
w
m
(
t
)
5:
c
1
,m
(
t
)
←
EMD(
x
m
(
t
))
▷
Apply
EMD
to
noisy
signal
6:
end
f
or
7:
c
1
(
t
)
←
1
M
P
M
m
=1
c
1
,m
(
t
)
▷
Ensemble
a
v
erage
of
rst
mode
8:
r
1
(
t
)
←
x
(
t
)
−
c
1
(
t
)
9:
f
or
i
=
2
to
n
do
10:
f
or
m
=
1
to
M
do
11:
r
i
−
1
,m
(
t
)
←
r
i
−
1
(
t
)
+
ε
·
w
m
(
t
)
12:
c
i,m
(
t
)
←
EMD(
r
i
−
1
,m
(
t
))
13:
end
f
or
14:
c
i
(
t
)
←
1
M
P
M
m
=1
c
i,m
(
t
)
▷
Ensemble
a
v
erage
of
i
-th
mode
15:
r
i
(
t
)
←
r
i
−
1
(
t
)
−
c
i
(
t
)
16:
end
f
or
17:
r
etur
n
{
c
1
(
t
)
,
.
.
.
,
c
n
(
t
)
}
and
r
n
(
t
)
CEEMD
AN
generate
a
series
of
ensembles
with
the
addition
of
controlled
white
noise
in
the
EMD
decomposition
operation.
This
helps
to
reduce
the
problem
of
mode
m
ixing
and
produces
smoother
and
Deep
hybrid
models
for
bitcoin
for
ecasting:
EMD,
CEEMD
AN,
and
LSTM
in
comparison
(A
youb
Aar
abi)
Evaluation Warning : The document was created with Spire.PDF for Python.
2802
❒
ISSN:
2252-8938
more
interpretable
IMFs.
Each
component
obtained
by
the
CEEMD
AN
decomposition
represents
independent
frequenc
y
characteristics
of
the
Bitcoin,
which
enhances
the
learning
capacity
of
the
LSTM
and
impro
v
es
the
rob
ustness
of
prediction.
The
ensemble-based
feature
of
CEEMD
AN
also
leads
to
better
signal
reconstruction
and
generalization
capabilities.
In
both
h
ybrid
methods,
IMF
components
(obtained
from
EMD
or
CEEMD
AN)
are
normalized
and
fed
into
separate
LSTM
netw
orks.
All
LSTMs
learn
the
dynamics
of
their
respecti
v
e
components.
Additi
v
e
predictions
for
all
IMFs
is
then
combined
to
get
the
nal
predicted
Bitcoin
price.
This
tw
o-step
process
separates
the
short-term
structures
(f
ast
IMFs)
from
the
long-term
structures
(slo
w
IMFs
or
the
residuals),
thus
impro
ving
the
accurac
y
of
the
model.
In
contrast,
CNN
architectures
such
as
those
used
for
series
or
image
classication
are
not
good
at
capturing
long
temporal
dependencies
and
do
not
generalize
well
to
v
ery
noisy
one-dimensional
signals.
Lik
e
wise,
classical
approaches
namely
,
ARIMA
and
GARCH
are
not
able
e
v
en
to
handle
stationary
series
in
such
a
non-stable
and
linear
re
gime
as
the
case
of
cryptocurrencies
[31]–[33].
3.5.
Gr
ey
r
elational
analysis
In
order
to
e
v
aluate
the
importance
of
the
f
actors
on
the
tar
get,
or
response
v
ariable
(price
of
Bitcoin),
the
so-called
GRA
introduced
in
the
gre
y
systems
theory
w
as
utilized
[34].
This
approach
allo
ws
one
to
estimate
the
le
v
el
of
similarity
of
time
series
(including
uncertainty
,
noise,
and
small
amounts
of
statistical
data).
The
reference
series,
corresponding
to
the
Bitcoin
price,
is
dened
as
(9).
X
0
=
{
x
0
(1)
,
x
0
(2)
,
...,
x
0
(
n
)
}
(9)
The
e
xplanatory
series
is
similarly
dened
in
(10).
X
i
=
{
x
i
(1)
,
x
i
(2)
,
...,
x
i
(
n
)
}
(10)
The
gre
y
relational
coef
cient
between
x
0
(
k
)
and
x
i
(
k
)
is
e
xpressed
by
(11).
ξ
i
(
k
)
=
min
i
min
k
|
x
0
(
k
)
−
x
i
(
k
)
|
+
ζ
·
max
i
max
k
|
x
0
(
k
)
−
x
i
(
k
)
|
|
x
0
(
k
)
−
x
i
(
k
)
|
+
ζ
·
max
i
max
k
|
x
0
(
k
)
−
x
i
(
k
)
|
(11)
Here,
ζ
∈
[0
,
1]
denotes
the
distinction
coef
cient,
which
is
usually
set
to
0.5.
The
o
v
erall
gre
y
relational
de
gree
between
X
0
and
X
i
is
then
computed
using
(12).
γ
i
=
1
n
n
X
k
=1
ξ
i
(
k
)
(12)
The
closer
γ
i
is
to
1,
the
higher
the
structural
similarity
between
the
e
xplanatory
v
ariabl
e
X
i
and
the
tar
get
series
X
0
.
This
method
is
well
adapted
to
non-stationary
nancial
time-series
such
as
those
e
xamined
in
[35]
and
to
the
problem
of
v
ariable
selection
prior
to
modeling.
3.6.
Ev
aluation
criteria
In
order
to
assess
the
forecasting
accurac
y
of
the
models
for
the
time
series
of
the
Bi
tcoin
price,
tw
o
typical
criteria
are
emplo
yed:
the
RMSE
and
the
mean
absolute
error
(MAE).
These
measurements
indicate
the
de
viation
between
predicted
and
actual
v
alues.
The
RMSE
metric
is
dened
in
(13).
R
M
S
E
=
v
u
u
t
1
n
n
X
t
=1
(
y
t
−
ˆ
y
t
)
2
(13)
It
represents
the
square
root
of
the
mean
squared
error
and
e
xpresses
the
a
v
erage
prediction
error
in
the
same
unit
as
the
data.
The
MAE
metric
is
dened
in
(14).
M
AE
=
1
n
n
X
t
=1
|
y
t
−
ˆ
y
t
|
(14)
MAE
measures
the
a
v
erage
absolute
prediction
error
and
is
generally
more
rob
ust
to
outliers
since
it
does
not
e
xcessi
v
ely
penalize
lar
ge
de
viations.
Int
J
Artif
Intell,
V
ol.
15,
No.
3,
June
2026:
2797–2810
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
2803
4.
D
A
T
A
COLLECTION
AND
SOURCES
The
purpose
of
this
study
is
to
forecast
the
path
of
the
Bitcoin
price
by
including,
be
yond
its
past
beha
vior
,
a
number
of
macro-nancial
v
ariables
that
can
potentially
af
fect
its
mo
v
ement
in
the
short/medium
term.
Because
the
cryptocurrenc
y
mark
et
is
di
v
erse,
as
control
v
ariables
were
selected
a
series
of
indicators
standing
for
the
global
economy
situat
ion,
the
U.S.
monetary
polic
y
,
the
capital
o
w
to
w
ards
safe-ha
v
en
assets,
and
the
dollar
dynamics.
Data
were
collected
o
v
er
a
ten-year
period
from
September
2014
to
September
2024,
with
a
daily
frequenc
y
and
standardized
prior
to
including
to
models.
4.1.
Crude
oil
price
W
est
T
exas
Intermediate
The
barometer
of
global
economic
acti
vity
that
is
the
oil
price
w
as
chosen
as
the
rst
e
xplanatory
v
ariable.
The
demand
for
oil
tends
to
rise
during
an
economic
e
xpansion,
and
f
all
when
the
economy
is
in
a
recession
or
is
slo
wing
do
wn
[36].
Additionally
,
the
v
olatility
of
oil
might
also
af
fect
the
e
xpectations
of
in
v
estors
re
g
arding
ination
and
monetary
polic
y
,
and
then
speculati
v
e
assets
such
as
Bitcoin
[37].
4.2.
Gold
price
(XA
U/USD)
Gold
has
long
been
considered
a
refuge
in
insecure
times
such
as
geopolitical
turm
oil
or
n
a
ncial
crises
[38].
Its
dynamics
typically
embody
the
anticipation
of
systemic
risk
and
ination.
When
the
mark
et
stresses,
in
v
estors
will
reduce
their
e
xposures
to
risk
y
assets,
and
the
risk
a
v
ersion
will
push
up
the
demand
for
gold;
the
latter
could
dri
v
e
capital
ight
out
of
cryptocurrencies
into
safer
assets
[39].
Therefore,
the
beha
viour
of
gold’
s
price
is
a
good
indicator
of
the
general
risk
a
v
ersion.
4.3.
2-y
ear
United
States
tr
easury
bond
rates
The
2-year
T
reasury
bond
yield
is
strongly
related
to
the
mark
et’
s
beliefs
about
the
monetary
polic
y
of
the
Fed
[40].
So
a
lo
w
rate
implies
an
accommodati
v
e
monetary
stance
(lo
wer
k
e
y
rates
and
e
xpanding
balance
sheet),
and
is
good
for
speculati
v
e
assets
and
the
stock
mark
et;
whereas
a
high
rate
speaks
of
a
monetary
tightening,
which
means
less
liquidity
and
more
stress
on
risk
y
assets.
In
this,
the
2Y
rate
is
a
k
e
y
dri
v
er
of
risk
appetite
perception
and
of
cross-asset
arbitrage
[41].
4.4.
United
States
dollar
index
The
dollar
inde
x
(DXY)
is
a
measure
of
the
v
alue
of
the
United
States
dollar
(U.S.
dollar)
relati
v
e
to
a
bask
et
of
six
foreign
currencies
(euro,
J
apanese
yen,
pound
sterling,
Canadian
dollar
,
Swedish
krona,
and
Swiss
fr
anc).
The
dollar
is
a
reserv
e
curre
nc
y
for
the
w
orld
and
is
frequentl
y
seen
as
a
s
afe
ha
v
en.
A
strong
U.S.
dollar
typically
means
a
do
wnturn
in
risk
appetite
or
increased
global
interest
rate
tightening,
while
a
weak
U.S.
dollar
can
mean
a
positi
v
e
en
vironment
for
non-U.S.
dollar
assets
[42].
Bitcoin,
at
times
touted
as
a
currenc
y-ination
hedge,
frequently
e
xperiences
a
ne
g
ati
v
e
dynamic
with
the
DXY
[43].
5.
RESUL
TS
AND
DISCUSSION
5.1.
Results
Figure
2
sho
ws
the
output
of
the
GRA
for
the
dif
ferent
e
xplanatory
v
ariables.
Each
of
the
lines
is
inde
x
ed
to
dif
ferent
time
interv
als:
daily
(1d),
daily
e
xcluding
week
ends
(1d
b),
weekly
(1w),
monthly
(1m),
and
quarterly
(1q)
series.
From
this
visualization,
it
is
possible
to
e
v
aluate
not
only
the
structure
of
relationship
on
the
basis
of
all
v
ariables
considered,
and
their
relation
to
the
price
of
Bitcoin,
b
ut
also
the
role
of
the
sampling
frequenc
y
on
the
capacity
to
e
xplain
it.
The
comparison
sho
ws
that
the
daily
data
generally
ha
v
e
the
highest
v
alues
of
GRA.
Crude
oil
price,
Gold
price,
US
2Y
interest
rate,
and
U.S.
dollar
inde
x
ha
v
e
a
close
GR
relationship
with
Bitcoin,
especially
in
high
frequenc
y
.
These
ndings
indicate
that
the
uctuations
of
these
indicators
are
highly
correlated
to
daily
changes
of
Bitcoin
trading.
Accordingly
,
the
model
w
as
designed
and
trained
using
only
the
da
ily
data,
as
it
is
the
scale
o
v
er
which
the
information
is
most
rele
v
ant
according
to
GRA.
This
option
can
capture
short-term
information
more
precisely
and
a
v
oid
informati
on
loss
c
aused
by
tem
p
or
al
aggre
g
ation.
It
also
mirrors
the
short
-term
horizon
upon
which
crypto
mark
ets
operate,
namely
their
acute
sensiti
vity
to
economic
data
and
geopolitical
shocks.
Figures
3
and
4
sho
w
the
e
xtracted
IMF
components
by
the
EMD
method
and
the
CEEMD
AN,
respecti
v
ely
.
It
is
noticed
that
the
separation
of
frequenc
y
presented
is
ner
and
less
noisy
in
CEEMD
AN,
resulting
in
better
-usable
signals
for
the
sequential
LSTM
training.
Especially
,
it
can
produ
c
e
a
bett
er
separation
of
high
frequenc
y
components,
at
the
same
time,
an
accurate
trend
is
more
distinct.
Deep
hybrid
models
for
bitcoin
for
ecasting:
EMD,
CEEMD
AN,
and
LSTM
in
comparison
(A
youb
Aar
abi)
Evaluation Warning : The document was created with Spire.PDF for Python.
2804
❒
ISSN:
2252-8938
Figure
2.
The
results
of
GRA
on
predictors
Figure
3.
The
IMFs
decomposed
by
EMD
Int
J
Artif
Intell,
V
ol.
15,
No.
3,
June
2026:
2797–2810
Evaluation Warning : The document was created with Spire.PDF for Python.
Int
J
Artif
Intell
ISSN:
2252-8938
❒
2805
Figure
4.
The
IMFs
decomposed
by
CEEMD
AN
The
performances
are
summarized
quantitati
v
ely
in
T
able
2.
The
proposed
baseline
LSTM
model
pro
vides
an
MAE
of
169.516
and
an
RMSE
of
256.225.
T
olerable
performances
can
be
achie
v
ed
by
including
EMD
decomposition
(MAE
=168.785;
RMSE
=256.042).
Last,
the
combination
of
CEEMD
AN
yi
elds
the
best
performance
(MAE
=167.837;
RMSE
=255.673)
o
v
erall.
Deep
hybrid
models
for
bitcoin
for
ecasting:
EMD,
CEEMD
AN,
and
LSTM
in
comparison
(A
youb
Aar
abi)
Evaluation Warning : The document was created with Spire.PDF for Python.
2806
❒
ISSN:
2252-8938
T
able
2.
Performance
comparison
between
LSTM,
EMD-LSTM,
and
CEEMD
AN-LSTM
models
Metric
LSTM
EMD-LSTM
CEEMD
AN-LSTM
MAE
169.516
168.785
167.837
RMSE
256.225
256.042
255.673
V
isual
predictions
are
visualized
in
Figures
5
to
7.
Figure
5
sho
ws
the
forecast
of
the
LSTM
model,
being
capable
to
recognize
the
global
trend,
it
misses
some
quick
re
v
ersals.
Figure
6
implemented
with
the
EMD-LSTM
model
possesses
more
sensiti
vity
to
local
change,
which
is
an
immediate
consequence
of
the
IMF
preprocessing.
Figure
7
sho
ws
that
CEEMD
AN-LSTM
presents
the
best
accurac
y
both
on
up
and
do
wn
trend
is
obtained
with
least
error
than
all
other
results.
The
learning
curv
es
de
v
elopment
indicates
that
the
CEEMD
AN
decomposition
is
capable
of
pro
viding
the
LSTM
model
more
solid
pattern
to
grasp
the,
thus
mitig
ating
the
inuence
of
abnormal
uctuations
of
the
Bitcoin
price.
Figure
5.
The
forecasting
results
based
on
LSTM
Figure
6.
The
forecasting
results
based
on
EMD-LSTM
Int
J
Artif
Intell,
V
ol.
15,
No.
3,
June
2026:
2797–2810
Evaluation Warning : The document was created with Spire.PDF for Python.