Structured Models¶
Introduction¶
VOR Stream models support two definition modes:
- Freeform — Write custom model code in Go, Python, or SAS. This gives full flexibility to implement any computation. See the Models page for details.
- Structured — Define a regression model declaratively by specifying predictors, coefficients, and a normalization method. No custom code is required.
Structured models are ideal when a model can be expressed as a regression equation. The UI guides you through building the formula, and the system computes the output at run time.
| Feature | Freeform | Structured |
|---|---|---|
| Definition method | Custom code (Go / Python / SAS) | Declarative (UI-driven) |
| Flexibility | Any computation | Regression equations |
| Requires coding | Yes | No |
| Normalization options | Implemented in code | Built-in (None, Logit, Probit) |
| Predictor management | Manual in code | UI table with coefficients and exponents |
Start here¶
Choose the page that matches what you need:
- New to structured models? Build your first structured model from sample input through a verified output.
- Applying a modeling pattern? Use the structured model use cases for categorical haircuts, regional factors, categorical scores and chained models.
- Looking up formula syntax? Go directly to the Formula Language reference.
Key concepts¶
- Regression Model
- The top-level structured definition. It specifies the function type, normalization method, and dependent variable (the target output).
- Regression Table
- Contains the intercept (baseline constant) and a list of numeric predictors that make up the regression equation.
- Numeric Predictor
- A single term in the regression equation. Each predictor has a coefficient, an exponent, and a field reference that identifies the data source.
- Field Reference
- Where a predictor gets its input value. A field reference is either direct, reading a value that already exists (Factor Reference, Dictionary Reference), or derived, computing one (Local Transformation, Lookup Reference). The two derived kinds are defined in the editor's Dynamic Fields panel. See Field references.
Mathematical formula¶
A structured regression model computes:
Y = f( β₀ + Σ(βᵢ × Xᵢ^eᵢ) )
where:
- Y is the dependent variable (model output)
- β₀ is the intercept
- βᵢ is the coefficient for predictor i
- Xᵢ is the value of predictor i
- eᵢ is the exponent for predictor i
- f() is the normalization function
Normalization methods¶
The normalization function f() transforms the raw linear combination into the final output:
| Method | Formula | Description |
|---|---|---|
| None | f(x) = x |
No transformation. Returns the raw linear combination. |
| Logit | f(x) = 1/(1 + e^(-x)) |
Logistic function. Maps output to the (0, 1) range. |
| Probit | f(x) = CDF(x) |
Cumulative standard normal distribution. Maps output to (0, 1). |
Logit and Probit are named for the model families they implement, and each applies that family's inverse link function: the logistic function for a logistic regression, the standard normal CDF for a probit regression. Choosing one here is how a structured model produces a probability, so there is nothing to compute in a formula.
Info
Use Logit or Probit normalization when the model output represents a probability, such as a probability of default. Use None when the raw linear combination is the desired output.
Field references¶
Every predictor draws on a field reference, which says where its input value comes from. There are four kinds, in two groups.
A direct field reference reads a value that already exists:
| Kind | Reads |
|---|---|
| Factor Reference | a risk factor supplied by the scenario |
| Dictionary Reference | a column of the portfolio or instrument data |
A derived field reference computes one from other field references:
| Kind | Produces its value by |
|---|---|
| Local Transformation | evaluating a formula |
| Lookup Reference | selecting a target based on a categorical value |
The Dynamic Fields panel is where the two derived kinds are defined, and that is the whole of what the term covers. Local Transformation and Lookup Reference are the two kinds defined there, so name the specific kind when it matters.
flowchart LR
D[Dictionary column] --> P[Predictor]
F[Scenario factor] --> P
D --> T[Local transformation]
F --> T
D --> L[Lookup reference]
F --> L
T --> P
T --> L
L --> P
P --> R[Structured regression]
R --> O[Dependent variable]
What each kind may reference¶
The combinations below are the supported dependency paths:
| Kind | Supported sources | Unsupported source |
|---|---|---|
| Predictor | any of the four kinds | the model's own output |
| Local Transformation | Dictionary References, Factor References | another Local Transformation, a Lookup |
| Lookup Reference | Dictionary References, Factor References, Local Transformations | another Lookup Reference |
The table has three practical consequences:
- Local Transformations accept Dictionary and Factor References. If two formulas need the same intermediate value, repeat the expression in both. Both are evaluated against the same inputs at the same point in the run, so the result is identical. Models submitted through the API or an import bypass the editor's sibling-reference check, save successfully and then fail at run time.
- Lookup References may target Local Transformations. Every formula is evaluated before the lookups resolve, so a lookup can use a transformation's result. The lookups themselves resolve in a single pass, which leaves a lookup naming another lookup unsupported. The editor does not offer one as a target, but a model arriving through the API or an import saves, and then resolves only when the inner lookup happens to run first.
- Model outputs are available to downstream nodes. To use a model's result as an input, chain a second model after it, as Estimate a PD, then stress it shows.
Factor reference¶
A factor reference links a predictor to a risk factor: an external or macroeconomic variable such as GDP growth, an unemployment rate, or a house price index.
Factors are managed in the Factors module and reach a model through the scenario attached to the study run. A factor derived from other factors is a Transformed Factor, which is evaluated once per scenario horizon. Use a Transformed Factor for calculations that need a time axis.
Dictionary reference¶
A dictionary reference links a predictor to a dictionary column: a field of the portfolio or instrument data, such as loan-to-value, credit score, debt-to-income ratio or account balance.
Dictionary columns are defined in the playpen's data dictionary and populated from the input data at run time.
Local transformation¶
A local transformation computes a value from a formula, evaluated once for each portfolio record. Use it to derive a predictor from the model's inputs, including a ratio between two columns or a log transform.
Write the formula in the formula language, which lists every operator and function available and covers conditionals, missing values and the rest of the syntax.
A local transformation is evaluated once per record, per scenario, per horizon. Repeated expressions add evaluation work on large scenario grids.
Name it distinctly from your data
Choose a local transformation name that is unique across the Dynamic Fields panel, input queue columns and factors. The editor permits collisions with input queue columns and factors, and the model then runs without error while one name carries two values. On an exact match the formula's result overwrites the data value, so every reader gets the computed value. When the names differ only in case both values survive under their own spellings: a predictor reads whichever spelling it names, and an output queue column of that name matches both spellings, so either value can land there, record by record. Either way the result is plausible and wrong for one of the name's two meanings.
The model reads input queue columns. An output-only column is safe and can be used for the audit technique in Apply a haircut by category.
Lookup reference¶
A lookup reference performs conditional field selection: it reads a categorical value and returns whichever field that category maps to. Use it when different segments of the portfolio should draw the same predictor from different sources.
Components:
- Selector: a categorical dictionary column, for example a province or state code.
- Mappings: a set of category-to-field pairs. Each names a category value and a target field reference: a factor, a dictionary column or a local transformation.
- Default Target: the field reference used when the selector value matches no mapping. Set one unless the mappings are certainly exhaustive; an uncovered selector value fails the run when the default is unset.
Name it distinctly from your data
Choose a lookup name that is unique across the Dynamic Fields panel, input queue columns and factors. The editor permits collisions with input queue columns and factors, which silently produce a plausible wrong result.
Switching between modes¶
A model can be switched between freeform and structured modes using the toggle button in the model editor.
Freeform to Structured: When switching from freeform to structured mode, freeform data (model script, series) remains saved and becomes inactive. The structured editor appears and you can define the regression. If you switch back to freeform, your previous script and series are still there.
Structured to Freeform: When switching from structured to freeform mode, the structured model definition is deleted when the model is saved. A warning dialog confirms this action before proceeding.
Note
Save the model to apply a mode switch. Until then, you can freely toggle between modes to compare them. Saving includes the active mode's data.
Next steps¶
- Build your first structured model.
- Apply a pattern from the structured model use cases.
- Look up operators and functions in the Formula Language.