The Best_Model flag identifies the selected forecast for
each series. Finn first checks candidate validity, then balances
backtest weighted MAPE with future-forecast plausibility. Individual
models, learned ensembles, and simple averages are evaluated on the same
series and expected backtest dates.
Model selection is several connected decisions, not one minimum-error calculation:
| Stage | Question answered | Result |
|---|---|---|
| Candidate validity | Does each model have complete, finite predictions and usable backtest accuracy? | Hard-invalid candidates cannot win or contribute to a simple average. |
| Within-run selection | Among sufficiently accurate candidates, which future path ranks best by risk, concern count, supported seasonal fidelity, WMAPE, and stable ID? | One selected model or exact average for each series. Soft concerns can remain when no better eligible alternative exists. |
| Agent iteration context | Which completed iteration’s settings should guide further optimization? | A near-best iteration can be preferred when its individual model-pool accuracy improves. This does not automatically replace a better saved local forecast. |
| Saved forecast promotion | Should the current iteration replace the saved result? | Local winners stay protected; globally selected series reference one winning global iteration. |
| Update acceptance | Is the newly refitted version of the saved choice still acceptable? | Keep that choice or attempt the existing default-local replacement workflow before reconciliation. |
Seasonal fidelity breaks a risk-and-concern tie only when every candidate in that tied group has a finite assessed score; otherwise WMAPE and stable ID decide the tie.
For a standard run, follow candidate checks, the accuracy allowance, and selection and saved averages. The Agent and update sections describe the additional decisions. Reconciliation makes the selected hierarchy coherent; it does not select the models again.
Two uses of “average” are important: a simple-average forecast combines predictions from selected components, while average model accuracy summarizes individual model WMAPEs to guide Agent search. Improving the latter can be useful even when the winning model has not changed.
These checks apply to candidate base forecasts before reconciliation,
including averages. Candidates with incomplete or non-finite predictions
are ineligible. Backtests and future predictions without a supported
trend are rejected when their absolute magnitude exceeds 100 times a
positive robust historical scale. That scale uses the 95th percentile of
absolute levels and robust level/change variation, not just the last
observation. A supported trend supplies a pointwise future bound of
100 * max(robust_scale, abs(projected_reference)), so
sustained growth is not rejected merely for accumulating over a long
horizon. The projection comes only from history, never from the
candidate being screened. Existing finite-negative handling still
follows negative_forecast, but missing or infinite
predictions are no longer replaced with zero.
Level and trend deviations beyond six robust reference scales create
soft concerns. The reference uses the latest available
max(12, 3 * seasonal_period, 2 * forecast_horizon) periods
before the historical cutoff. Without supported drift, its path is
seasonal naive when a complete seasonal cycle is available, otherwise a
recent median level. The robust scale is the largest of the 95th
percentile of absolute levels, level MAD, and one-step-change MAD. In
that fallback, level-reference width is the larger of the relevant
change MAD and 5% of that scale, multiplied by
6 * sqrt(horizon_step) for the comparison. Supported drift
instead supplies a projected center and a width reflecting residual
noise and slope variation, described below. Trend checks compare median
changes over matched historical and forecast spans, with a robust scale
floor; proportional references perform those comparisons on log
changes.
With at least two complete finite historical cycles, a detrended seasonal profile can support additional checks. Seasonal strength must be at least 0.6 and the period at least three. For forecasts covering at least one cycle, negative profile correlation or amplitude outside one-third to three times historical amplitude creates a concern. Supported proportional trends use log-scale historical and forecast profiles, so constant proportional seasonality is not confused with growing absolute amplitude. Additive and fallback paths retain the original seasonal calculations. Unsupported checks remain unassessed. These are guardrails, not calibrated prediction intervals or guarantees of future accuracy.
For a shorter horizon with at least three future points, Finn compares only the upcoming phases. Their historical range must exceed both three times the historical phase-residual MAD and 5% of the full historical seasonal amplitude. Finn removes the robust historical deseasonalized trend from the future path, rather than fitting a free trend through those few points. Negative correlation creates a phase concern only when future variation also exceeds that noise threshold. Full-cycle amplitude concern bounds are not applied to a partial cycle.
One or two seasonal cycles remain usable. One cycle supplies a cautious reference, while repeated cycles support stronger comparisons. Weak seasonality, fewer than three future points, low-variation partial phases, and unavailable checks remain unassessed instead of making the series fail. Flat forecasts are not rejected simply because observations are noisy.
| Check | Trigger | Effect | When Unassessed |
|---|---|---|---|
| Coverage | Missing, duplicate, or unexpected expected scenario/date keys in backtests or future predictions | Hard rejection | Never for a candidate being selected |
| Finite predictions | Any required NA, NaN, Inf, or
-Inf prediction |
Hard rejection | Never |
| Catastrophic magnitude | Absolute backtest or unsupported-trend future prediction exceeds
100 * robust_scale; supported future step exceeds
100 * max(robust_scale, abs(projected_reference))
|
Hard rejection | When historical scale is zero |
| Accuracy availability | No finite weighted MAPE from usable actuals | Hard rejection | Never |
| Level | Forecast departs from the supported trend reference, or seasonal-naive/recent-median fallback, beyond its horizon-adjusted width | Soft concern | Skipped after a hard failure |
| Zero-history level | A nonzero future forecast when the reference history is entirely zero | Soft unsupported_level concern |
Zero forecasts pass; this is not a ratio-based hard rejection |
| Trend | Median forecast change differs from historical matched-span changes beyond six robust slope scales, using log changes for a supported proportional reference | Soft concern | Fewer than four matched periods, non-finite reference values, incompatible log-domain forecasts, or a hard failure |
| Seasonal amplitude | Future detrended amplitude is below one-third or above three times the historical profile | Soft concern | Insufficient cycles, strength below 0.6, period below three, short forecast, or zero historical profile amplitude |
| Seasonal phase | Future and historical profiles have negative correlation after the appropriate trend adjustment | Soft concern | Insufficient historical cycles or strength, period below three, undefined correlation, or insufficient points or signal for a partial horizon |
| Seasonal amplitude preference | Amplitude difference exceeds the tolerance learned from historical cycle variation | Tie-break only; no additional concern or rejection | Insufficient historical evidence, zero reference amplitude, or insufficient partial-horizon phase signal |
For these checks the default seasonal period is daily 7,
weekly 52, monthly 12, quarterly
4, and yearly 1. If multiple seasonal periods
are configured, the smallest valid integer greater than one is used. The
checks do not independently validate every seasonal pattern a model can
learn. The trend scale is the larger of historical matched-span slope
MAD and 0.05 * working_scale / matched_span. The working
scale is the original normalized robust scale for additive or fallback
paths, and the analogous robust scale of normalized log values for
proportional paths.
Each assessed soft check has a normalized excess score. Candidate risk is the maximum of the level, trend, and seasonality scores, not their average. The concern count counts the triggered reasons; seasonal phase and amplitude can contribute two reasons. Unassessed checks do not count as concerns.
The separate in-memory Seasonal_Fidelity score measures
amplitude distortion beyond historical variation. For the phases being
assessed, let reference_amplitude be the reference profile
range, future_amplitude the comparable future range, and
cycle_amplitudes the ranges of centered historical cycle
profiles. The tolerance is
max(3 * MAD(cycle_amplitudes), 0.05 * reference_amplitude).
Fidelity is the finite nonnegative score
max(0, (abs(future_amplitude - reference_amplitude) - tolerance) / reference_amplitude).
Insufficient evidence gives NA, not an artificial zero. A
partial-horizon flat average can have an assessed amplitude preference
even when its phase correlation is undefined. This score does not change
risk or the concern count, and by itself cannot reject reuse, prevent
accuracy-goal stopping, or trigger another fit.
A trend reference requires at least
max(12, 3 * seasonal_period) finite, regularly spaced
prepared-history values in the existing reference window. Missing dates
are not bridged, and missing values are not dropped to manufacture
regular history. If explicit observation flags supplied to the evaluator
identify missing or unobserved fitting values, trend support is
declined. Existing prepared artifacts do not establish such provenance
automatically; retained imputation is still part of their evidence.
Finn compares at most two robust reference forms, without training any additional forecasting model:
log(target) - log(normalization). It is considered only
when all historical values are strictly positive and exceed
sqrt(.Machine$double.eps) * robust_scale. Signed, zero, and
relatively near-zero history is not shifted, clipped, or transformed
with log1p to make this form fit.With period p, two chronological prefixes end at
n - 2q and n - q, where
q = max(1, floor(p / 2)). Each predicts the next
q historical values and must contain at least
max(8, 2p) training values. Drift must be nonzero, have the
same sign in both prefixes and the full window, and exceed twice the MAD
of its per-step slope estimates, with a numerical-precision floor. Each
validation block must improve MAE by at least 20% over the no-drift
reference on the same dates. Comparisons occur on the original target
scale, normalized by a common factor for numerical stability.
Proportional drift displaces an accepted additive reference only if it
improves additive MAE by at least 20% in both blocks. Zero-error and
negligible comparisons retain the simpler reference. These fixed support
thresholds are engineering choices, not statistical confidence
statements.
For a supported reference, let sigma_error be the larger
of its residual MAD and a noise floor, and sigma_slope the
MAD of its per-step drift estimates. The additive floor is 5% of the
existing normalized robust scale; the log-space floor is
log1p(0.05). At step h, the level-comparison
width is:
Overlapping slope estimates are not treated as independent observations to shrink uncertainty. The projected center follows the supported drift and upcoming seasonal phase; it does not widen itself in response to candidate forecasts. Reference estimation is shared across candidates through the existing evaluation cache. Nonrepresentable reference projections fall back to the original checks instead of clipping forecasts or removing the magnitude bound.
For example, a positive series with well-supported 2% monthly growth can have a reference near 126.8 after 12 months and 160.8 after 24 months when its current fitted level is 100 and it has no seasonal effect. Continuing that growth need not produce a level or absolute-slope penalty merely because the horizon is longer. A path that accelerates beyond the supported trajectory can still receive a concern or fail its pointwise magnitude bound. This reference is not substituted for the selected model’s actual forecast.
A finite zero or negative forecast under a supported proportional
reference receives a level_deviation soft concern, while
log-only slope and seasonal comparisons are unassessed. The forecast is
not changed, and this does not introduce a new hard sign veto. Short,
irregular, missing, or unstable histories and all-zero series retain the
earlier fallback behavior. New scoring applies to newly evaluated
predictions; a complete saved winner is still reused without
retrospective reassessment. Unexpected regimes and mixed seasonal
mechanisms can remain uncertain, and an eligible singleton with soft
concerns can still win under the ordinary best-available policy.
After hard failures are excluded, let best_wmape be the
smallest eligible weighted MAPE as a fraction. Candidates within
best_wmape + max(0.005, 0.05 * best_wmape) form the
shortlist. Selection prefers lower forecast risk, then fewer concerns.
Within each risk-and-concern tie, smaller seasonal amplitude distortion
precedes WMAPE only when every tied candidate has a finite assessed
fidelity score. Otherwise that tie retains the existing WMAPE ordering.
A stable candidate identifier breaks the final tie. Missing seasonal
evidence is not rewarded as zero distortion.
For example, with a best eligible WMAPE of 8.0%, the ceiling is 8.5%. A safer 8.3% candidate can beat a questionable 8.0% candidate, but an 8.8% candidate is outside the allowance. At 20% the ceiling is 21%; at 2% it is 2.5%. The absolute allowance is relatively generous for very accurate series and is an initial engineering default, not a calibrated statistical bound.
For ordinary and iterative forecasts, when every shortlisted eligible candidate has soft concerns, Finn selects the best available one and reports them. When all candidates hard-fail, Finn raises an error. Update reuse and default replacements have a stricter acceptance boundary, described below. No policy can guarantee that a good forecast exists for every input.
forecast_time_series() finishes through
final_models(). A manual standard workflow uses the same
process when final_models() is called:
max_model_average. A non-finite
component is not concealed by averaging with
na.rm = TRUE.Best_Model = "Yes". If it is
an average, save that exact combination in the existing average-model
artifact. Otherwise rank just the eligible averages by the same policy
and save the winner of that subset with Best_Model = "No".
This is not necessarily the lowest-WMAPE average. If no eligible average
can be formed, there is no required average artifact.hts solver and
publish its bottom-level result. Do not run another future-quality
selection after reconciliation. Format intervals and any weekly-to-daily
allocation using the existing output workflow.Individual outputs remain available with their best-model flags, but only the selected simple average is retained, not every computed combination. Standard forecasting selects among models already run; a quality failure does not fit an additional model automatically.
On retry, Finn validates the combined saved outputs before reusing a
winner. A Best_Model column, an average filename, or a
finite completion log alone is insufficient. All individual models
marked "No" can be correct when a complete saved average is
marked "Yes". Missing, incomplete, or ambiguous selection
flags cause averages and selection to be rebuilt from existing
predictions, without training again. Saved weekly outputs are restored
to native cadence before recomputation. Complete winners are not
re-evaluated for future quality, and a multi-series retry retains both
previously completed and newly finalized results. A saved average whose
original component outputs are unavailable requires restoration of those
artifacts; it is not silently reconstructed from a different subset.
Quality-aware selection happens before reconciliation. Different
hierarchy nodes can choose different models or averages. Finn reconciles
that selected mixture with the existing hts::combinef()
nodes/groups, residual weights, and nonnegative settings, then publishes
the resulting bottom-level forecasts.
There is no runtime plausibility scoring of reconciled future paths, no switch of the whole output to a uniform reconciled model family, and no replacement of individual bottom rows after the solve. Standard runs may still save per-model reconciled outputs for inspection, but those outputs are not automatically promoted by a post-reconciliation quality ranker. An outer Agent reconciliation solves its selected mixture directly.
Reconciled backtests provide normal bottom-level WMAPE reporting and run comparison. The solver consumes selected forecast values, hierarchy structure, and residual-based weights, not quality rankings. A resumed run does not require every source node’s prior ranking or repeat its future-quality assessment. Actual complete selected source forecasts are still required when a solve must run. Weekly source forecasts are evaluated at weekly cadence during selection, and a non-finite value on any daily-expanded forecast row remains non-finite when its weekly key is restored.
Reconciliation can dampen, spread, or amplify a problem. Better base forecasts can improve bottom-level output, but passing the source checks is not a guarantee that every reconciled path is plausible. Solver and ordinary artifact errors remain errors. Test-only audits can measure the resulting forecasts without changing production selection.
iterate_forecast() runs final_models()
within each submitted forecast run. Iteration ranking starts with the
earliest minimum-WMAPE result in the current Agent version. Later
eligible iterations within 10% relative of that WMAPE can become the
preferred search context when their model_avg_wmape is
strictly lower; the lowest such average wins, with earlier ties
retained. Missing or non-finite averages do not create an improvement.
This is separate from the within-iteration 0.5-percentage-point/5%
allowance: risk, concern count, and seasonal fidelity are not reassessed
to rank past iterations.
For local iterations, model_avg_wmape,
model_median_wmape, and model_std_wmape
summarize the individual model candidates’ backtest WMAPEs, excluding
simple-average forecasts. For example, ARIMA may remain best after
adding an external regressor while the multivariate models improve. A
lower average preserves that promising search direction because further
iterations may make those models the winners; it is not a guarantee of
future accuracy. Global summary fields retain their original meanings:
overall run WMAPE for mean and median, and zero spread. Statistics use
already-loaded current backtests; history comparisons reuse recorded
metrics without reloading past forecasts. Unavailable local statistics
are not fabricated from the selected model’s score.
Choosing an iteration’s settings for further optimization is distinct
from replacing a saved forecast: a better saved local forecast remains
protected. Global promotion is decided once for the complete iteration,
then all saved global winners move together even if one series
individually worsens. All globally selected series must reference one
best_run_name; they may still use different models or
averages from that iteration. A partial global evaluation cannot promote
only its successful series. Reload, finalization, and update reject
mixed global iteration metadata, including inconsistent state left by
interrupted writes, rather than silently using multiple global
iterations. These per-series writes are not a storage transaction. Agent
comparisons, saved WMAPE, and accuracy-goal decisions retain
four-decimal precision.
Normal accuracy-goal stopping and local optimization routing use completeness and WMAPE, without another soft-quality veto. Consequently, an iteration winner with soft concerns can beat an earlier winner or meet the accuracy goal; hard-invalid and incomplete results still cannot claim successful completion. Rejected runs consume iteration budget. At the limit, an eligible best-available result may be retained with concerns; if none exists, Finn fails or uses an already enabled local phase for unresolved global series. Final outer reconciliation does not initiate another quality-selection or refitting loop. Avoided evaluator calls and artifact reads are covered by tests; no fixed runtime reduction is promised.
For ordinary Agent iterations, the goal and saved-forecast comparisons use selected completed-output backtests, including daily-expanded rows when weekly forecasts are delivered daily. Within-run candidate selection stays at native cadence. Those scores are not interchangeable: the zero-actual convention and rounding can produce different WMAPEs after daily allocation. Future target placeholders never contribute to either accuracy calculation.
update_forecast() follows a different fitting path but
uses the same evaluator:
final_models() to search over a new candidate pool.max(10, ceiling(0.20 * existing_series)).final_models() for each default replacement, then
require the replacement to pass all applicable quality checks.
Short-history checks that cannot be assessed do not count as failures.
Good backtest accuracy alone cannot publish a still-questionable
replacement.The optional accuracy-degradation allow_iterate_forecast
workflow remains separate. Checks after refit or retune assess newly
generated predictions, not old iteration winners. A default’s saved
acceptance or rejection prevents repeating that acceptance decision
after it has completed; an interrupted acceptance step can still assess
the new output once. Quality recovery itself does not ask the LLM to
waive a check. Default refitting can increase compute or
configured-provider cost, and it is a bounded recovery attempt rather
than a promise of a satisfactory forecast.
Update accuracy keeps the existing native-cadence aggregate used for refit/retune and aggregate logging. Completed-output backtests supply per-series saved-winner accuracy and local model-pool statistics. This distinction preserves the update workflow’s metric meanings when weekly output is expanded to daily rows.
Developer tests compare two paths on the same fake candidate outputs:
accuracy-only selection followed by real hts, and
quality-aware selection/averaging followed by real hts.
Held-out future truth is used only by the tests, never supplied to the
selector. They measure bottom-level error and shape, including effects
on uninjected siblings. Small regressions run with package tests;
broader stress and full-catalogue averaging timings are separate
developer experiments.
A separate bounded, in-memory corner-test matrix checks exact prediction keys, signed/zero/missing-actual accuracy, threshold boundaries, seasonal alignment, optional-score ties, weekly group isolation, and replacement-selector contracts. Fixed-seed cases also check independent accuracy and ranking expectations, row-order invariance, and the inability of a hard-invalid candidate to displace a valid winner. These tests do not fit forecast models or run additional reconciliation, and passing them is not proof that every possible input or future regime change is handled correctly.
Earlier experiments exposed half-amplitude averages and short-horizon phase reversals that the original concern checks did not distinguish. Focused regressions now prefer an intact available seasonal candidate within the accuracy allowance and assess informative partial horizons using repeated historical evidence. Controls retain naturally varying seasonal amplitudes within the history-adaptive tolerance. These results do not calibrate a guarantee: one-cycle histories, fewer than three future points, or low-signal partial phases can still retain reversed timing because assessment is unsupported. A plausible alternative outside the WMAPE allowance cannot win a soft-risk or fidelity comparison. Unexpected regime changes are not known in advance, and the existing nonnegative solver floor can produce small positive values for all-zero histories. None of these limits invokes a post-reconciliation replacement model.
The number of simple averages grows as
choose(N, 2) + choose(N, 3) when
max_model_average = 3. The normal smoke test uses four
mixed model/recipe candidates and independently checks all ten
pair/triple averages. The separate developer timing workload uses every
supported model/recipe output and reports averaging and evaluation
separately, with fresh artifacts for every trial. Earlier timing and
broad stress measurements predate the refined seasonal ranking; they are
not new measurements of this revision. No new elapsed-time assertion or
production timeout is added. Runtime depends on candidate count,
horizon, backtests, storage, and the machine; measurements should
precede any future regression limit.
Checks use the existing prepared data for each series, preferring R1
when prepared and otherwise normalizing R2’s Horizon == 1
rows. Differencing and Box-Cox are reversed using existing metadata.
Target_Original is used when available so outlier cleaning
does not change the actuals used for evaluation; otherwise
Target is used. Prior missing-value imputation is retained.
Recipe and metadata files are read by exact path, without directory
listing, and history is reused in memory across candidates.
Selection requires at least one candidate and at least one finite actual inside the recent reference window. An empty candidate pool or a reference window containing only missing/non-finite actuals raises a clear input error. Finite observations farther back, or future target values, cannot silently supply that window’s evidence. Partly missing windows remain usable when a finite value, including zero, is present; unavailable seasonal checks remain unassessed, and missing backtest actuals continue to receive no accuracy weight.
Original and cleaned targets have separate inverse-differencing starting values. Older second-order original-target artifacts missing those values must be regenerated from original input; already saved forecast outputs remain readable.
Rankings and explanations are held in memory. No additional diagnostic or reference files are written. The selector has an internal replaceable function contract; a public per-run custom-evaluator API is not included in this release.
A pointwise absolute percentage error is produced first, rounded to
four decimal places, then weighted by the absolute actual value across
expected backtest rows. The existing convention replaces zero actuals
with 0.1 for this calculation. Original actuals are matched
by date; missing/non-finite actuals do not contribute weight, and a
candidate without a finite accuracy score is ineligible. Overlapping
backtests keep their separate scenario/horizon observations. The
positive-target example below illustrates the weighting.
#> Simple Back Test Results
#> # A tibble: 10 × 8
#> Combo Date Model FCST Target MAPE Target_Total Percent_Total
#> <chr> <date> <chr> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 Country_1 2020-01-01 arima 9 10 0.1 150 0.0667
#> 2 Country_1 2020-02-01 arima 23 20 0.15 150 0.133
#> 3 Country_1 2020-03-01 arima 35 30 0.167 150 0.2
#> 4 Country_1 2020-04-01 arima 41 40 0.025 150 0.267
#> 5 Country_1 2020-05-01 arima 48 50 0.04 150 0.333
#> 6 Country_1 2020-01-01 ets 7 10 0.3 150 0.0667
#> 7 Country_1 2020-02-01 ets 22 20 0.1 150 0.133
#> 8 Country_1 2020-03-01 ets 29 30 0.0333 150 0.2
#> 9 Country_1 2020-04-01 ets 42 40 0.05 150 0.267
#> 10 Country_1 2020-05-01 ets 53 50 0.06 150 0.333
#>
#> Overall Model Accuracy by Combo
#> # A tibble: 2 × 4
#> Combo Model MAPE Weighted_MAPE
#> <chr> <chr> <dbl> <dbl>
#> 1 Country_1 arima 0.0963 0.0800
#> 2 Country_1 ets 0.109 0.0733
During the simple back test process above, arima seems to be the better model from a pure MAPE perspective, but ETS ends up being the winner when using weighted MAPE. The benefits of weighted MAPE allow finnts to find the optimal model that performs the best on the biggest components of a forecast, which comes with the added benefit of putting more weight on more recent observations since those are more likely to have larger target values then ones further into the past. Another way of putting more weight on more recent observations is how Finn overlaps its back testing scenarios. This means the most recent observations are tested for accuracy in different forecast horizons (H=1, H=2, etc). More info on this in the back testing vignette.
Users can evaluate the retained model outputs with their own metrics and choose another model. The internal selector is replaceable, but this release does not expose a public per-run evaluator or persist custom selection functions.