Select Best Models and Prep Final Outputs
final_models(
run_info,
average_models = TRUE,
max_model_average = 3,
weekly_to_daily = TRUE,
parallel_processing = NULL,
inner_parallel = FALSE,
num_cores = NULL
)run info using the set_run_info() function.
If TRUE, create simple averages of individual models and save the eligible average selected by accuracy and future-quality checks.
Max number of models to average together. Will create model averages for 2 models up until input value or max number of models ran.
If TRUE, convert a week forecast down to day by evenly splitting across each day of week. Helps when aggregating up to higher temporal levels like month or quarter.
Default of NULL runs no parallel processing and forecasts each individual time series one after another. 'local_machine' leverages all cores on current machine Finn is running on. 'spark' runs time series in parallel on a spark cluster in Azure Databricks or Azure Synapse.
Run components of forecast process inside a specific time series in parallel. Can only be used if parallel_processing is set to NULL or 'spark'.
Number of cores to run when parallel processing is set up. Used when running parallel computations on local machine or within Azure. Default of NULL uses total amount of cores on machine minus one. Can't be greater than number of cores on machine minus 1.
Final model outputs are written to disk.
Candidates are screened for complete, finite predictions and extreme
future magnitudes using original-scale prepared actuals. Among eligible
candidates within 0.5 percentage points or 5 percent relative weighted MAPE
of the best eligible accuracy, whichever allowance is larger, selection
prefers lower risk and fewer future level, trend, and seasonal concerns.
Tied candidates with assessed seasonal evidence then prefer smaller amplitude
distortion beyond historical cycle variation, before weighted MAPE and model
identifier. This preference alone does not reject a forecast. Repeated strong
historical seasonality can also support phase checks over at least three
informative future periods, even when the horizon is shorter than a cycle.
Short histories remain usable; unsupported seasonal checks are not assessed.
Sufficient regular prepared history can support an additive or proportional
trend reference when it improves both chronological historical validation
blocks. Level checks then follow the projected trend and seasonal offsets,
with uncertainty from historical residuals and drift variation. Proportional
references use log-scale changes and seasonal comparisons only for suitable
positive history. Unsupported or unstable trends retain the existing
seasonal-naive or recent-median reference. Backtests retain the historical
magnitude bound; supported future paths use 100 times the larger of the
historical robust scale and the absolute projected reference at each step.
Prepared-history imputation remains part of the evidence. These checks are
engineering guardrails, not calibrated intervals or accuracy guarantees.
If all candidates fail the required checks, selection raises an error.
Evaluation is deterministic for fixed inputs and creates no diagnostic files.
If an individual model wins, the best eligible simple average is still saved
with Best_Model = "No", using the same quality-aware ranking among averages.
If an average wins overall, that exact average is saved as the best model.
No average artifact is required when no eligible average can be formed.
Quality selection happens before hierarchical reconciliation, at each prepared
hierarchy node. The selected mixture is reconciled without a second future
plausibility evaluation or a whole-hierarchy replacement model. Reconciled
backtests still supply reported accuracy; reconciliation does not require
retained quality rankings. On retry, saved individual and average outputs
must identify one complete winner per series. A Best_Model column or an
average filename alone is not proof of completion. Incomplete selections
rebuild averages and winner flags from existing predictions without fitting
models again. Complete saved winners are reused without future-quality
reassessment, and every series remains in the returned result. Reconciled
output is reused only when all original series have complete, unique backtest
and future keys with finite forecasts, including every day of a daily-expanded
week. Incomplete reconciled output is rebuilt from selected source forecasts
and checked before completion is logged; storage and read errors propagate.
During model-result repair of an accepted default forecast, old average
outputs and selection-completion flags are not reused. Selection is rebuilt
from the replacement individual predictions through the same policy, with
stale averages reset using the ordinary forecast schema. This recovery does
not change preparation or allow repeated genuinely rejected defaults.
If only one individual remains eligible, the schema-correct empty optional
average represents no average and remains readable during publication and
restart. Empty required artifacts and malformed optional averages are errors.
# \donttest{
data_tbl <- timetk::m4_monthly %>%
dplyr::rename(Date = date) %>%
dplyr::mutate(id = as.character(id)) %>%
dplyr::filter(
Date >= "2013-01-01",
Date <= "2015-06-01"
)
run_info <- set_run_info()
#> Finn Submission Info
#> • Project Name: finn_project
#> • Run Name: finn_fcst-20261002T154256Z
#>
prep_data(run_info,
input_data = data_tbl,
combo_variables = c("id"),
target_variable = "value",
date_type = "month",
forecast_horizon = 3
)
#> ℹ Prepping Data
#> ✔ Prepping Data [2.9s]
#>
prep_models(run_info,
models_to_run = c("arima", "ets"),
back_test_scenarios = 3
)
#> ℹ Creating Model Workflows
#> ✔ Creating Model Workflows [139ms]
#>
#> ℹ Creating Model Hyperparameters
#> ✔ Creating Model Hyperparameters [119ms]
#>
#> ℹ Creating Train Test Splits
#> ℹ Turning ensemble models off since no multivariate models were chosen to run.
#> ℹ Creating Train Test Splits
#> ✔ Creating Train Test Splits [284ms]
#>
train_models(run_info,
run_global_models = FALSE
)
#> ℹ Training Individual Models
#> ✔ Training Individual Models [24.6s]
#>
final_models(run_info)
#> ℹ Selecting Best Models
#> ✔ Selecting Best Models [856ms]
#>
# }