The AI agent is a tool-calling orchestration layer that sits on top of the core finnts pipeline. It uses an LLM to:
You keep control through a few inputs (data, horizon, optional regressors, performance goal, iteration budget); the agent does the rest.
Core Agent Functions:
iterate_forecast(): Run the agent to iterate toward a
best forecast run.update_forecast(): Update forecasts with new data,
using models trained in previous agent runs, optionally re-invoking the
agent if accuracy degrades.ask_agent(): Ask natural language questions about your
forecast results and get data-driven answers.agent_version and run_id) and log files.Use these helpers to retrieve outputs:
get_best_agent_run(agent_info)get_agent_forecast(agent_info)ask_agent(agent_info, question)finntsSet up environment variables for Azure OpenAI (example):
Sys.setenv(
AZURE_OPENAI_ENDPOINT = "<your-endpoint>",
AZURE_OPENAI_API_KEY = "<your-key>",
AZURE_OPENAI_API_VERSION = "<api-version>"
)TIMESFM_API_URL and TIMESFM_API_TOKEN). It is
a local-only model that does not support external regressors.
Below is a complete flow using the built-in M4 monthly sample.
library(finnts)
library(dplyr)
project <- set_project_info(
project_name = "ai_agent_demo",
path = tempdir(), # or a persistent folder
combo_variables = c("id"),
target_variable = "value",
date_type = "month", # day|week|month|quarter|year
fiscal_year_start = 1 # fiscal month (1 = Jan)
)Tip:
pathcontrols where logs/forecasts/EDA artifacts are saved.
Supports local filesystem, Azure Blob (viaAzureStor::blob_container), or Microsoft 365 drives (ms365r) viastorage_object.
Date (class
Date).target_variable.
hist_data <- timetk::m4_monthly %>%
dplyr::filter(date >= as.Date("2013-01-01")) %>%
dplyr::rename(Date = date) %>%
dplyr::mutate(id = as.character(id))llm is an ellmer Chat that Finn uses as a configuration
template. Finn never mutates this template. It creates an isolated,
empty-history session for each forecast series and each
ask_agent() request, so one modern model can perform both
forecast input selection and results analysis without conversation
history leaking between workflows.
llm <- ellmer::chat_azure_openai(model = "gpt-4o-mini")
agent <- set_agent_info(
project_info = project,
llm = llm,
input_data = hist_data,
forecast_horizon = 6, # number of future periods
external_regressors = NULL, # e.g., c("Price","Promo")
allow_hierarchical_forecast = FALSE, # set TRUE to let agent use hierarchies
negative_forecast = FALSE, # set TRUE to allow forecasts below zero
overwrite = TRUE # start a fresh run_id if inputs changed
)This writes the versioned inputs into path/input_data/
(hashed by combo/run) and logs the new
agent_version/run_id.
iterate_forecast(
agent_info = agent,
weighted_mape_goal = 0.05, # your accuracy target of 5%
max_iter = 3, # stop after N iterations if not hitting goal
)What happens under the hood:
final_models() selects
the winner within each iteration. Iteration ranking starts with the
earliest minimum WMAPE and can prefer a later result within 10% relative
when its average model WMAPE is lower. This preserves improvements
across other models even when the current winner is unchanged, without
repeating past future-quality checks.max_iter. A soft warning does not
impose another stopping veto, while hard-invalid or incomplete results
cannot count as success.Reconciliation uses the selected forecasts and backtest residuals, not retained quality rankings. Incomplete selection retries rebuild averages and winner flags from existing predictions without retraining. All globally selected series share one winning global iteration, although their chosen models or averages can differ within it. Global updates preserve those saved subsets from that one run and check the newly produced forecasts before acceptance; they do not reintroduce every model from the original request or combine different global iterations.
best_runs <- get_best_agent_run(agent_info = agent, full_run_info = TRUE)
head(best_runs)
fcst <- get_agent_forecast(agent_info = agent)
head(fcst)best_runs summarizes, for each time series combo, the
best run inputs when calling the Finn forecast process.fcst returns the consolidated forecast table. For
hierarchical Agent runs, it contains only the final reconciled
Best-Model forecast. Series can run and select different
models, recipes, or averages during iterate_forecast(), so
there is no complete per-model comparison hierarchy to publish. The best
forecast is the reconciled combination of selected series forecasts, not
necessarily one identical model family for every series. This best-only
reconciled output also applies after update_forecast();
non-hierarchical candidate output is unchanged.For non-agentic hierarchical runs, use
get_forecast_data(run_info) to retrieve every successfully
saved per-model reconciled forecast plus Best-Model. Filter
Best_Model == "Yes" when only the selected forecast is
needed.
After running iterate_forecast() or
update_forecast(), you can use ask_agent() to
ask natural language questions about your results. The agent analyzes
your forecast data, model configurations, and EDA outputs to provide
data-driven answers.
ask_agent() creates an LLM-driven workflow that: 1.
Plans the analysis steps needed to answer your question
2. Executes R code to analyze the relevant data 3.
Generates a natural language answer based on the
results
# Ask about forecast accuracy
answer <- ask_agent(
agent_info = agent,
question = "What is the average weighted MAPE across all time series?"
)
# Ask about models used
answer <- ask_agent(
agent_info = agent,
question = "Which models were selected as best for each time series?"
)
# Ask about feature importance
answer <- ask_agent(
agent_info = agent,
question = "What are the top 3 most important features for the forecast models?"
)
# Ask about data quality
answer <- ask_agent(
agent_info = agent,
question = "Were there any missing values or outliers in the data?"
)
# Ask about specific forecasts
answer <- ask_agent(
agent_info = agent,
question = "What are the forecasted values for M750 for the next 3 months?"
)
# Ask about time series characteristics
answer <- ask_agent(
agent_info = agent,
question = "Which time series show strong seasonality patterns?"
)
# Ask comparative questions
answer <- ask_agent(
agent_info = agent,
question = "Which time series have the highest forecast uncertainty?"
)ask_agent() has access to four main data sources:
get_agent_forecast()): Future predictions, back-test
results, model selections, confidence intervalsget_best_agent_run()): Feature engineering settings,
transformations applied, model hyperparametersget_eda_data()): Time
series characteristics, seasonality, stationarity tests, data quality
metricsget_summarized_models()): Feature importance, model
parameters, recipe detailsThe agent automatically determines which data sources to use based on your question.
When you have new input data, keep the same project
and create a new agent run with updated
input_data. Then call update_forecast():
# suppose you've appended more months to hist_data:
hist_data2 <- hist_data %>% dplyr::filter(Date <= as.Date("2016-06-01"))
agent2 <- set_agent_info(
project_info = project,
llm = llm,
input_data = hist_data2,
forecast_horizon = 6,
overwrite = TRUE # required to create a new agent version when running update_forecast()
)
update_forecast(
agent_info = agent2,
weighted_mape_goal = 0.05,
allow_iterate_forecast = TRUE, # if degradation detected, allow the agent to re-iterate
max_iter = 2 # cap re-iteration cost
)
updated_fcst <- get_agent_forecast(agent2)
# Ask questions about the updated forecast
answer <- ask_agent(
agent_info = agent2,
question = "Summarize the forecast accuracy."
)What update_forecast() does:
Validation or Ensemble in the saved
splits are excluded only from delivery-completion checks, for both
winners and average components; all saved rows remain intact. Unknown
scenario IDs still invalidate completion. Valid results are skipped;
missing or damaged results are scheduled for the ordinary refit and
overwritten at their existing paths. Shared global results are handled
together, without refitting separately for each series or replacing
valid local winners. Known recipe settings and worker splits are reused,
and required context is read by exact path. Series without current
best-run metadata follow the normal update path without an additional
driver artifact audit.iterate_forecast() to retrain from
scratch.final_models(). It invokes the
shared evaluator after refitting and after any retuning, before a
hierarchy is reconciled. Required components must pass hard eligibility
and their selected combination must have no applicable future-quality
concerns. An incomplete or rejected reused hierarchy is not solved; its
covered current series follow the default-local path. Quality-rejected
current series receive one default reforecast independently of the
ordinary execution-failure limit. The default run uses
final_models() and must pass both hard eligibility and
applicable soft checks before success. Reforecasting many rejected
series can increase runtime and provider cost.allow_iterate_forecast = TRUE, it will
invoke the iterate loop (bounded by
max_iter) to recover accuracy.hts solver produces bottom-level forecasts from the
accepted selected source rows. Finn does not evaluate reconciled future
outputs to replace the hierarchy with a uniform model family or trigger
a late default-refitting loop. Solver and artifact errors remain errors;
passing source quality checks is not a guarantee of future
accuracy.Within each Agent iteration, final_models() selects
candidates using accuracy and future-quality checks. The earliest
minimum-WMAPE iteration anchors the comparison; a later eligible result
within 10% relative can supply the next search context when its average
model WMAPE is strictly lower. Local mean, median, and standard
deviation describe the individual-model backtests, excluding simple
averages; global summaries retain run WMAPE for mean and median and zero
spread. ARIMA can therefore remain the best model while improved
multivariate models after an xreg change preserve a promising search
direction. That does not overwrite a better saved local forecast. Global
promotion moves all global winners to one iteration together. Partial
evaluations or interrupted writes cannot silently publish a mixed global
selection. Normal goal stopping uses complete eligible results and
four-decimal WMAPE, without a second soft-quality veto. Hierarchical
comparisons use reconciled backtests, and recorded metrics avoid
reassessing past future paths. Rejected evaluations still consume
iteration budget. Newly generated update forecasts and default
replacements retain the stricter acceptance checks described above.
Within an iteration’s accuracy allowance, candidates tied on risk and concern count can prefer smaller seasonal-amplitude distortion beyond historical cycle variation. That preference applies only when every tied candidate has an assessed score; missing evidence falls back to WMAPE. For hierarchical candidates, this preference applies at each source node before reconciliation; it does not rank completed Agent runs. Fidelity alone is not a quality rejection, a reason to keep iterating past the accuracy goal, or a trigger for default refitting. Repeated strong historical seasonality can support phase checks for informative horizons of at least three points even when they are shorter than a cycle. Unsupported seasonal checks remain unassessed.
The shared evaluator can follow historically supported additive or proportional growth across long horizons. Two chronological historical comparisons must support the drift before it replaces the seasonal-naive or recent-median reference. Level uncertainty includes residual and slope variation; proportional paths use compatible log-scale trend and seasonal checks. Supported future magnitudes are screened against the larger of historical scale and the projected reference at each step, while backtest and fallback bounds remain unchanged. This uses existing prepared history, including retained imputation, and does not add an observation-provenance guarantee, a wider accuracy allowance, or another fitting loop. Complete saved winners are not retroactively rescored.
Selection uses original actuals from exact prepared-data reads and keeps detailed rankings in memory. Native weekly and daily-expanded saved source forecasts reconstruct the same selection evidence; invalid values on later expanded days cannot disappear during weekly restoration. Only existing run bookkeeping records selected runs, evaluated/rejected status, and default-recovery acceptance for restart safety. No new quality-log files or serialized evaluation functions are created.
See Best Model Selection for the exact accuracy allowance, hard and soft checks, saved-average behavior, and the separate standard, iterative, and update workflows. Ordinary and iterative best-available soft-concern behavior is not the stricter acceptance rule used for update reuse and default replacements.
allow_hierarchical_forecast controls which series enter
Agent optimization:
FALSE keeps the original bottom-level series. Global
iterations can still test the exact standard or
grouped hierarchy detected by EDA; Finn reconciles that
candidate to the bottom level before comparing WMAPE. Local iterations
remain bottoms_up.TRUE detects the hierarchy and expands the input to all
hierarchy levels before optimization. Those prepared levels use a single
ID combo column, so all global and local iterations use
bottoms_up; Finn then performs one final reconciliation
using the detected hierarchy and publishes the bottom-level result
without post-reconciliation quality selection.The detected structure can be none, standard (for
example, Region → Country → SKU), or grouped (crossed
dimensions). In either mode, hierarchical candidates are compared at the
bottom level and the final get_agent_forecast() output
contains bottom-level forecasts.
For background and manual control, see the “Hierarchical Forecasting” vignette.
If you pass
external_regressors = c("Price","Promo", ...):
input_data for the selected
columns.parallel_processing = "local_machine" runs each time
series in parallel across local cores.parallel_processing = "spark" executes combos on an
Azure Databricks/Synapse Spark cluster (see “Parallel
Processing” vignette).inner_parallel = TRUE parallelizes work
inside a combo (useful when outer parallelism is
NULL or "spark").num_cores = NULL defaults to all cores minus
one.Every time-series combo receives independent driver and reasoning
Chat objects with empty conversation history, whether execution is
sequential or parallel. Parallel runs require ellmer 0.4.0 or later on
the driver and every worker; Finn serializes the configured Chats
through foreach before creating the per-combo deep
clones.
You normally won’t need this, but for audits:
path/input_data/…
path/eda/…
path/logs/… (includes the hashed
*-agent_run.csv and *-agent_best_run.* for
each version)path/final_output/…
Use the helpers first; dig into files only if you must.