Selects an approach compatible with FinnTS hierarchy construction without changing input data, running preprocessing, or saving artifacts.

detect_hierarchy(
  input_data,
  combo_variables,
  target_variable = NULL,
  combo_cleanup_date = NULL,
  hist_end_date = NULL
)

Arguments

input_data

A local data frame containing the combo columns. Repeated observations are allowed. Collect Spark data frames before calling.

combo_variables

Nonempty character vector of unique combo column names.

target_variable

Target column name, required only when combo_cleanup_date is supplied. Must contain numeric values; missing targets are allowed but infinite values are not.

combo_cleanup_date

Optional Date scalar. When supplied, apply the same inactive-series cleanup as prep_data(): retain series whose target sum between this date and hist_end_date, inclusive, is nonzero.

hist_end_date

Date scalar required for cleanup. Future target values do not contribute to the cleanup sum.

Value

A list with forecast_approach ("standard_hierarchy", "grouped_hierarchy", or "bottoms_up"), hierarchy_order, named distinct counts, retained bottom-series count total_ts, a conflicts data frame (parent, child, child_label, parent_count), and a readable reason. Diagnostics are returned in memory only.

Details

Character and factor combo boundaries are trimmed using the same normalization as prep_data(). The caller's data is unchanged. Missing, blank, nonfinite, or unsupported combo labels, ambiguous "--"-joined combo identities, invalid cleanup inputs, and an empty retained population produce errors.

Levels are ordered by increasing distinct count, with supplied column order breaking ties, matching the engine. A standard hierarchy requires every child label to determine exactly one parent at each adjacent level and the finest level to identify every bottom-level tuple. Crossed dimensions or reused child labels select a grouped hierarchy. One retained series selects "bottoms_up" because the HTS constructors require multivariate input.

Detection describes observed relationships, not unobserved business relationships. Use the same input population and cleanup settings that will be passed to prep_data(). This function does not override explicit preprocessing choices. Agent setup shares the structural analysis but preserves its single-column bottoms-up policy and saved-version contracts.

Examples

data <- data.frame(
  Region = c("North", "North", "South"),
  Site = c("A", "B", "C")
)
decision <- detect_hierarchy(data, c("Region", "Site"))
decision$forecast_approach
#> [1] "standard_hierarchy"

data$Date <- as.Date("2026-08-01")
data$Revenue <- c(10, 0, 20)
detect_hierarchy(
  data, c("Region", "Site"), target_variable = "Revenue",
  combo_cleanup_date = as.Date("2025-08-01"),
  hist_end_date = as.Date("2026-08-01")
)
#> $forecast_approach
#> [1] "standard_hierarchy"
#> 
#> $hierarchy_order
#> [1] "Region" "Site"  
#> 
#> $counts
#> Region   Site 
#>      2      2 
#> 
#> $total_ts
#> [1] 2
#> 
#> $conflicts
#> [1] parent       child        child_label  parent_count
#> <0 rows> (or 0-length row.names)
#> 
#> $reason
#> [1] "Every child has one parent in engine order."
#>