Why an AI agent in finnts?

The AI agent is a tool-calling orchestration layer that sits on top of the core finnts pipeline. It uses an LLM to:

  • run lightweight EDA (missing data, outliers, (P)ACF, seasonality, stationarity),
  • select/adjust feature recipes (lags, rolling windows, Fourier, date feats),
  • choose model families (local vs global; ARIMA/ETS/LM/Tree/GBM/NN, etc.), backtests, and ensembling,
  • optionally detect and use hierarchies (bottoms-up, standard, grouped),
  • generate a reproducible best run and a consolidated forecast artifact you can pull with a single call.

You keep control through a few inputs (data, horizon, optional regressors, performance goal, iteration budget); the agent does the rest.

Core Agent Functions:

  • iterate_forecast(): Run the agent to iterate toward a best forecast run.
  • update_forecast(): Update forecasts with new data, using models trained in previous agent runs, optionally re-invoking the agent if accuracy degrades.
  • ask_agent(): Ask natural language questions about your forecast results and get data-driven answers.

What the agent produces

  • A versioned agent run (with agent_version and run_id) and log files.
  • A best-run table per time series with metrics.
  • A combined forecast table (optionally reconciled if hierarchy is used).
  • EDA artifacts (missing data, outliers, stationarity, seasonality, ACF/PACF).
  • Summaries of trained models (hyperparameters, recipes, feature importance).

Use these helpers to retrieve outputs:

  • get_best_agent_run(agent_info)
  • get_agent_forecast(agent_info)
  • ask_agent(agent_info, question)

Prerequisites

  • Package: finnts
  • LLM driver: the agent uses an ellmer Chat object for tool calling.
    • Example here shows Azure OpenAI, but any ellmer Chat backend works.

Set up environment variables for Azure OpenAI (example):

Sys.setenv(
  AZURE_OPENAI_ENDPOINT    = "<your-endpoint>",
  AZURE_OPENAI_API_KEY     = "<your-key>",
  AZURE_OPENAI_API_VERSION = "<api-version>"
)
  • TimeGPT Setup: In order to use TimeGPT model as part of agent forecast workflow, set the TimeGPT credentials.
  • Chronos2 Setup: In order to use Chronos2 model as part of agent forecast workflow, set the Chronos2 API credentials.
  • Chronos Bolt Base Setup: Uses the same API credentials as Chronos2. It is a lighter-weight model that does not support external regressors and runs as a local model only.
  • Chronos Bolt Tiny Setup: Uses the same API credentials as Chronos2/Chronos Bolt Base. It is the smallest/fastest model variant that does not support external regressors and runs as a local model only.
  • TimesFM Setup: Uses its own API endpoint (TIMESFM_API_URL and TIMESFM_API_TOKEN). It is a local-only model that does not support external regressors.

End-to-end: first run with the AI agent

Below is a complete flow using the built-in M4 monthly sample.

1) Create a project

library(finnts)
library(dplyr)

project <- set_project_info(
  project_name = "ai_agent_demo",
  path = tempdir(), # or a persistent folder
  combo_variables = c("id"),
  target_variable = "value",
  date_type = "month", # day|week|month|quarter|year
  fiscal_year_start = 1 # fiscal month (1 = Jan)
)

Tip: path controls where logs/forecasts/EDA artifacts are saved.
Supports local filesystem, Azure Blob (via AzureStor::blob_container), or Microsoft 365 drives (ms365r) via storage_object.

2) Bring data

  • One row per time series combo/date.
  • A single Date column named Date (class Date).
  • Target column matches target_variable.
  • Optional external regressors as extra columns (historical or historical+future values, if you want them used in forecasting).
hist_data <- timetk::m4_monthly %>%
  dplyr::filter(date >= as.Date("2013-01-01")) %>%
  dplyr::rename(Date = date) %>%
  dplyr::mutate(id = as.character(id))

3) Define the LLM

llm is an ellmer Chat that Finn uses as a configuration template. Finn never mutates this template. It creates an isolated, empty-history session for each forecast series and each ask_agent() request, so one modern model can perform both forecast input selection and results analysis without conversation history leaking between workflows.

llm <- ellmer::chat_azure_openai(model = "gpt-4o-mini")

4) Create the agent run

agent <- set_agent_info(
  project_info = project,
  llm = llm,
  input_data = hist_data,
  forecast_horizon = 6, # number of future periods
  external_regressors = NULL, # e.g., c("Price","Promo")
  allow_hierarchical_forecast = FALSE, # set TRUE to let agent use hierarchies
  negative_forecast = FALSE, # set TRUE to allow forecasts below zero
  overwrite = TRUE # start a fresh run_id if inputs changed
)

This writes the versioned inputs into path/input_data/ (hashed by combo/run) and logs the new agent_version/run_id.

5) Let the agent iterate to a best run

iterate_forecast(
  agent_info          = agent,
  weighted_mape_goal  = 0.05, # your accuracy target of 5%
  max_iter            = 3, # stop after N iterations if not hitting goal
)

What happens under the hood:

  • Reads cached EDA (or computes it) via the agent’s EDA tools.
  • Chooses seasonal period, missing/outlier handling, box-cox/differencing strategy.
  • Sweeps models (local &/or global), backtests, recipes (lags/rolling/Fourier/date feats).
  • Optionally enables external regressors if they improve WMAPE.
  • Evaluates future quality when final_models() selects the winner within each iteration. Iteration ranking starts with the earliest minimum WMAPE and can prefer a later result within 10% relative when its average model WMAPE is lower. This preserves improvements across other models even when the current winner is unchanged, without repeating past future-quality checks.
  • Stops early when a complete eligible winner meets the WMAPE goal; otherwise iterates up to max_iter. A soft warning does not impose another stopping veto, while hard-invalid or incomplete results cannot count as success.

Reconciliation uses the selected forecasts and backtest residuals, not retained quality rankings. Incomplete selection retries rebuild averages and winner flags from existing predictions without retraining. All globally selected series share one winning global iteration, although their chosen models or averages can differ within it. Global updates preserve those saved subsets from that one run and check the newly produced forecasts before acceptance; they do not reintroduce every model from the original request or combine different global iterations.

6) Retrieve results

best_runs <- get_best_agent_run(agent_info = agent, full_run_info = TRUE)
head(best_runs)

fcst <- get_agent_forecast(agent_info = agent)
head(fcst)
  • best_runs summarizes, for each time series combo, the best run inputs when calling the Finn forecast process.
  • fcst returns the consolidated forecast table. For hierarchical Agent runs, it contains only the final reconciled Best-Model forecast. Series can run and select different models, recipes, or averages during iterate_forecast(), so there is no complete per-model comparison hierarchy to publish. The best forecast is the reconciled combination of selected series forecasts, not necessarily one identical model family for every series. This best-only reconciled output also applies after update_forecast(); non-hierarchical candidate output is unchanged.

For non-agentic hierarchical runs, use get_forecast_data(run_info) to retrieve every successfully saved per-model reconciled forecast plus Best-Model. Filter Best_Model == "Yes" when only the selected forecast is needed.


Ask questions about your forecast results

After running iterate_forecast() or update_forecast(), you can use ask_agent() to ask natural language questions about your results. The agent analyzes your forecast data, model configurations, and EDA outputs to provide data-driven answers.

How it works

ask_agent() creates an LLM-driven workflow that: 1. Plans the analysis steps needed to answer your question 2. Executes R code to analyze the relevant data 3. Generates a natural language answer based on the results

Example questions

# Ask about forecast accuracy
answer <- ask_agent(
  agent_info = agent,
  question = "What is the average weighted MAPE across all time series?"
)

# Ask about models used
answer <- ask_agent(
  agent_info = agent,
  question = "Which models were selected as best for each time series?"
)

# Ask about feature importance
answer <- ask_agent(
  agent_info = agent,
  question = "What are the top 3 most important features for the forecast models?"
)

# Ask about data quality
answer <- ask_agent(
  agent_info = agent,
  question = "Were there any missing values or outliers in the data?"
)

# Ask about specific forecasts
answer <- ask_agent(
  agent_info = agent,
  question = "What are the forecasted values for M750 for the next 3 months?"
)

# Ask about time series characteristics
answer <- ask_agent(
  agent_info = agent,
  question = "Which time series show strong seasonality patterns?"
)

# Ask comparative questions
answer <- ask_agent(
  agent_info = agent,
  question = "Which time series have the highest forecast uncertainty?"
)

What data sources are available

ask_agent() has access to four main data sources:

  1. Forecast results (get_agent_forecast()): Future predictions, back-test results, model selections, confidence intervals
  2. Model configurations (get_best_agent_run()): Feature engineering settings, transformations applied, model hyperparameters
  3. EDA results (get_eda_data()): Time series characteristics, seasonality, stationarity tests, data quality metrics
  4. Model summaries (get_summarized_models()): Feature importance, model parameters, recipe details

The agent automatically determines which data sources to use based on your question.

Tips for effective questions

  • Be specific about what you want to know
  • Reference specific time series by their combo ID if needed
  • Ask about metrics, patterns, or comparisons
  • The agent can perform calculations and aggregations on the fly

Updating with new data (production loop)

When you have new input data, keep the same project and create a new agent run with updated input_data. Then call update_forecast():

# suppose you've appended more months to hist_data:
hist_data2 <- hist_data %>% dplyr::filter(Date <= as.Date("2016-06-01"))

agent2 <- set_agent_info(
  project_info = project,
  llm = llm,
  input_data = hist_data2,
  forecast_horizon = 6,
  overwrite = TRUE # required to create a new agent version when running update_forecast()
)

update_forecast(
  agent_info             = agent2,
  weighted_mape_goal     = 0.05,
  allow_iterate_forecast = TRUE, # if degradation detected, allow the agent to re-iterate
  max_iter               = 2 # cap re-iteration cost
)

updated_fcst <- get_agent_forecast(agent2)

# Ask questions about the updated forecast
answer <- ask_agent(
  agent_info = agent2,
  question = "Summarize the forecast accuracy."
)

What update_forecast() does:

  • Rebuilds global models first (if used), requiring all global winners to reference one iteration, then updates local series that need it. Mixed global iteration metadata fails before refitting instead of dispatching several global updates.
  • Verifies recorded completions: On restart, the driver reads the selected fitted models and required forecasts for series that have current best-run metadata. Completion requires every expected date and backtest scenario from the saved preparation calendar and splits, including every day of expanded weekly forecasts. Rows declared Validation or Ensemble in the saved splits are excluded only from delivery-completion checks, for both winners and average components; all saved rows remain intact. Unknown scenario IDs still invalidate completion. Valid results are skipped; missing or damaged results are scheduled for the ordinary refit and overwritten at their existing paths. Shared global results are handled together, without refitting separately for each series or replacing valid local winners. Known recipe settings and worker splits are reused, and required context is read by exact path. Series without current best-run metadata follow the normal update path without an additional driver artifact audit.
  • Avoids repeated fitting inside workers: Before refitting, each worker checks whether another attempt already completed the current models and forecasts. Complete results can be reused, including finishing missing best-run logging when existing run settings and metrics are sufficient. This does not repeat quality selection or date conversion. CSV run, series, and model identifiers retain their exact text. A repaired default with one eligible model can publish without an average; only a schema-correct empty optional average is accepted, not malformed or required empty artifacts. Only the existing combo task list and worker context are dispatched; there are no additional saved fields or preparation-repair controls. Required preparation and predecessor metadata access failures remain hard errors rather than default-model fallbacks. These checks do not lock files or prevent an active duplicate from overwriting them later.
  • Handles new time series: If new series appear in the data (up to 20% of existing series, floor of 10), simple forecasts are created automatically using default local model inputs—no LLM involvement needed.
  • Handles failed time series: If individual time series fail during the global or local model update (e.g., due to data issues or model errors), they are automatically re-forecast using the same default local model inputs as new series. If more than 20% of existing series (floor of 10) fail to update, the run errors out and directs you to use iterate_forecast() to retrain from scratch.
  • Checks reused forecasts before reconciliation: The reuse path does not call final_models(). It invokes the shared evaluator after refitting and after any retuning, before a hierarchy is reconciled. Required components must pass hard eligibility and their selected combination must have no applicable future-quality concerns. An incomplete or rejected reused hierarchy is not solved; its covered current series follow the default-local path. Quality-rejected current series receive one default reforecast independently of the ordinary execution-failure limit. The default run uses final_models() and must pass both hard eligibility and applicable soft checks before success. Reforecasting many rejected series can increase runtime and provider cost.
  • Compares WMAPE to a trailing baseline of previous runs. If >40% of series are >20% worse than the previous run WMAPE, and allow_iterate_forecast = TRUE, it will invoke the iterate loop (bounded by max_iter) to recover accuracy.
  • Reconciles the selected mixture: The existing hts solver produces bottom-level forecasts from the accepted selected source rows. Finn does not evaluate reconciled future outputs to replace the hierarchy with a uniform model family or trigger a late default-refitting loop. Solver and artifact errors remain errors; passing source quality checks is not a guarantee of future accuracy.

Within each Agent iteration, final_models() selects candidates using accuracy and future-quality checks. The earliest minimum-WMAPE iteration anchors the comparison; a later eligible result within 10% relative can supply the next search context when its average model WMAPE is strictly lower. Local mean, median, and standard deviation describe the individual-model backtests, excluding simple averages; global summaries retain run WMAPE for mean and median and zero spread. ARIMA can therefore remain the best model while improved multivariate models after an xreg change preserve a promising search direction. That does not overwrite a better saved local forecast. Global promotion moves all global winners to one iteration together. Partial evaluations or interrupted writes cannot silently publish a mixed global selection. Normal goal stopping uses complete eligible results and four-decimal WMAPE, without a second soft-quality veto. Hierarchical comparisons use reconciled backtests, and recorded metrics avoid reassessing past future paths. Rejected evaluations still consume iteration budget. Newly generated update forecasts and default replacements retain the stricter acceptance checks described above.

Within an iteration’s accuracy allowance, candidates tied on risk and concern count can prefer smaller seasonal-amplitude distortion beyond historical cycle variation. That preference applies only when every tied candidate has an assessed score; missing evidence falls back to WMAPE. For hierarchical candidates, this preference applies at each source node before reconciliation; it does not rank completed Agent runs. Fidelity alone is not a quality rejection, a reason to keep iterating past the accuracy goal, or a trigger for default refitting. Repeated strong historical seasonality can support phase checks for informative horizons of at least three points even when they are shorter than a cycle. Unsupported seasonal checks remain unassessed.

The shared evaluator can follow historically supported additive or proportional growth across long horizons. Two chronological historical comparisons must support the drift before it replaces the seasonal-naive or recent-median reference. Level uncertainty includes residual and slope variation; proportional paths use compatible log-scale trend and seasonal checks. Supported future magnitudes are screened against the larger of historical scale and the projected reference at each step, while backtest and fallback bounds remain unchanged. This uses existing prepared history, including retained imputation, and does not add an observation-provenance guarantee, a wider accuracy allowance, or another fitting loop. Complete saved winners are not retroactively rescored.

Selection uses original actuals from exact prepared-data reads and keeps detailed rankings in memory. Native weekly and daily-expanded saved source forecasts reconstruct the same selection evidence; invalid values on later expanded days cannot disappear during weekly restoration. Only existing run bookkeeping records selected runs, evaluated/rejected status, and default-recovery acceptance for restart safety. No new quality-log files or serialized evaluation functions are created.

See Best Model Selection for the exact accuracy allowance, hard and soft checks, saved-average behavior, and the separate standard, iterative, and update workflows. Ordinary and iterative best-available soft-concern behavior is not the stricter acceptance rule used for update reuse and default replacements.


Hierarchies (optional)

allow_hierarchical_forecast controls which series enter Agent optimization:

  • FALSE keeps the original bottom-level series. Global iterations can still test the exact standard or grouped hierarchy detected by EDA; Finn reconciles that candidate to the bottom level before comparing WMAPE. Local iterations remain bottoms_up.
  • TRUE detects the hierarchy and expands the input to all hierarchy levels before optimization. Those prepared levels use a single ID combo column, so all global and local iterations use bottoms_up; Finn then performs one final reconciliation using the detected hierarchy and publishes the bottom-level result without post-reconciliation quality selection.

The detected structure can be none, standard (for example, Region → Country → SKU), or grouped (crossed dimensions). In either mode, hierarchical candidates are compared at the bottom level and the final get_agent_forecast() output contains bottom-level forecasts.

For background and manual control, see the “Hierarchical Forecasting” vignette.


External regressors (xregs)

If you pass external_regressors = c("Price","Promo", ...):

  • Provide historical or historical+future values (length ≥ horizon) in your input_data for the selected columns.
  • The agent will test which specific xregs help; if they don’t, it will disable them for that iteration.
  • See the “External Regressors” vignette for data prep details.

Parallelism knobs

  • parallel_processing = "local_machine" runs each time series in parallel across local cores.
  • parallel_processing = "spark" executes combos on an Azure Databricks/Synapse Spark cluster (see “Parallel Processing” vignette).
  • inner_parallel = TRUE parallelizes work inside a combo (useful when outer parallelism is NULL or "spark").
  • num_cores = NULL defaults to all cores minus one.

Every time-series combo receives independent driver and reasoning Chat objects with empty conversation history, whether execution is sequential or parallel. Parallel runs require ellmer 0.4.0 or later on the driver and every worker; Finn serializes the configured Chats through foreach before creating the per-combo deep clones.


Reading artifacts directly (optional)

You normally won’t need this, but for audits:

  • Inputs: path/input_data/…
  • EDA: path/eda/…
  • Logs: path/logs/… (includes the hashed *-agent_run.csv and *-agent_best_run.* for each version)
  • Final Agent Outputs: path/final_output/…

Use the helpers first; dig into files only if you must.