主要内容

armaxOptions

R2026b

Option set for armax

Description

Use an armaxOptions object to specify options for estimating parameters of ARMAX, ARIMAX, ARMA or ARIMA models through the armax function. You can specify options such as the handling of initial states or the ability to generate parameter covariance data.

Creation

Description

opt = armaxOptions creates the default options set for armax.

example

opt = armaxOptions(Name,Value) creates an option set with the options specified by one or more Name,Value pair arguments.

example

Properties

expand all

Handling of initial conditions during estimation, specified as one of the following values:

  • 'zero' — The initial conditions are set to zero.

  • 'estimate' — The initial conditions are treated as independent estimation parameters.

  • 'backcast' — The initial conditions are estimated using the best least squares fit.

  • 'auto' — The software chooses the method to handle initial conditions based on the estimation data.

Error to be minimized in the loss function during estimation, specified as the comma-separated pair consisting of 'Focus' and one of the following values:

  • 'prediction' — The one-step ahead prediction error between measured and predicted outputs is minimized during estimation. As a result, the estimation focuses on producing a good predictor model.

  • 'simulation' — The simulation error between measured and simulated outputs is minimized during estimation. As a result, the estimation focuses on making a good fit for simulation of model response with the current inputs.

The Focus option can be interpreted as a weighting filter in the loss function. For more information, see Loss Function and Model Quality Metrics.

Weighting prefilter applied to the loss function to be minimized during estimation. To understand the effect of WeightingFilter on the loss function, see Loss Function and Model Quality Metrics.

Specify WeightingFilter as one of the following values:

  • [] — No weighting prefilter is used.

  • Passbands — Specify a row vector or matrix containing frequency values that define desired passbands. You select a frequency band where the fit between estimated model and estimation data is optimized. For example, [wl,wh], where wl and wh represent lower and upper limits of a passband. For a matrix with several rows defining frequency passbands, [w1l,w1h;w2l,w2h;w3l,w3h;...], the estimation algorithm uses the union of the frequency ranges to define the estimation passband.

    Passbands are expressed in rad/TimeUnit for time-domain data and in FrequencyUnit for frequency-domain data, where TimeUnit and FrequencyUnit are the time and frequency units of the estimation data.

  • SISO filter — Specify a single-input-single-output (SISO) linear filter in one of the following ways:

    • A SISO LTI model

    • {A,B,C,D} format, which specifies the state-space matrices of a filter with the same sample time as estimation data.

    • {numerator,denominator} format, which specifies the numerator and denominator of the filter as a transfer function with same sample time as estimation data.

      This option calculates the weighting function as a product of the filter and the input spectrum to estimate the transfer function.

Control whether to enforce stability of estimated model, specified as the comma-separated pair consisting of 'EnforceStability' and either true or false.

This option is not available for multi-output models with a non-diagonal A polynomial array.

Data Types: logical

Option to generate parameter covariance data, specified as true or false.

If EstimateCovariance is true, then use getcov to fetch the covariance matrix from the estimated model.

Option to display the estimation progress, specified as one of the following values:

  • 'on' — Information on model structure and estimation results are displayed in a progress-viewer window.

  • 'off' — No progress or results information is displayed.

Removal of offset from time-domain input data during estimation, specified as one of the following:

  • A column vector of positive integers of length Nu, where Nu is the number of inputs.

  • [] — Indicates no offset.

  • Nu-by-Ne matrix — For multi-experiment data, specify InputOffset as an Nu-by-Ne matrix. Nu is the number of inputs and Ne is the number of experiments.

Each entry specified by InputOffset is subtracted from the corresponding input data.

Removal of offset from time-domain output data during estimation, specified as one of the following:

  • A column vector of length Ny, where Ny is the number of outputs.

  • [] — Indicates no offset.

  • Ny-by-Ne matrix — For multi-experiment data, specify OutputOffset as a Ny-by-Ne matrix. Ny is the number of outputs, and Ne is the number of experiments.

Each entry specified by OutputOffset is subtracted from the corresponding output data.

Options for regularized estimation of model parameters, specified as a structure with the fields in the following table. For more information on regularization, see Regularized Estimates of Model Parameters.

Field NameDescriptionDefault
Lambda

Constant that determines the bias versus variance tradeoff.

Specify a positive scalar to add the regularization term to the estimation cost.

The default value of 0 implies no regularization.

0
R

Weighting matrix.

Specify a vector of nonnegative numbers or a square positive semi-definite matrix. The length must be equal to the number of free parameters of the model.

For black-box models, using the default value is recommended. For structured and grey-box models, you can also specify a vector of np positive numbers such that each entry denotes the confidence in the value of the associated parameter.

The default value of 1 implies a value of eye(npfree), where npfree is the number of free parameters.

1
Nominal

The nominal value towards which the free parameters are pulled during estimation.

The default value of 0 implies that the parameter values are pulled towards zero. If you are refining a model, you can set the value to 'model' to pull the parameters towards the parameter values of the initial model. The initial parameter values must be finite for this setting to work.

0

Numerical search method used for iterative parameter estimation, specified as the one of the values in the following table.

SearchMethodDescription
'auto'

Automatic method selection

A combination of the line search algorithms, 'gn', 'lm', 'gna', and 'grad', is tried in sequence at each iteration. The first descent direction leading to a reduction in estimation cost is used.

'gn'

Subspace Gauss-Newton least-squares search

Singular values of the Jacobian matrix less than GnPinvConstant*eps*max(size(J))*norm(J) are discarded when computing the search direction. J is the Jacobian matrix. The Hessian matrix is approximated as JTJ. If this direction shows no improvement, the function tries the gradient direction.

'gna'

Adaptive subspace Gauss-Newton search

Eigenvalues less than gamma*max(sv) of the Hessian are ignored, where sv contains the singular values of the Hessian. The Gauss-Newton direction is computed in the remaining subspace. gamma has the initial value InitialGnaTolerance (see Advanced in 'SearchOptions' for more information). This value is increased by the factor LMStep each time the search fails to find a lower value of the criterion in fewer than five bisections. This value is decreased by the factor 2*LMStep each time a search is successful without any bisections.

'lm'

Levenberg-Marquardt least squares search

Each parameter value is -pinv(H+d*I)*grad from the previous value. H is the Hessian, I is the identity matrix, and grad is the gradient. d is a number that is increased until a lower value of the criterion is found.

This algorithm requires Optimization Toolbox™ software.

'grad'

Steepest descent least-squares search

'lsqnonlin'

Trust-region-reflective algorithm of lsqnonlin (Optimization Toolbox)

This algorithm requires Optimization Toolbox software.

'patternsearch'

Solver for nonlinearities without well-defined gradients

You can use the patternsearch (Global Optimization Toolbox) solver to find the minimum of a nonlinear function that does not have a well-defined gradient. This solver requires Global Optimization Toolbox software.

'fmincon'

Constrained nonlinear solvers

You can use the sequential quadratic programming (SQP) and trust-region-reflective algorithms of the fmincon (Optimization Toolbox) solver. If you have Optimization Toolbox software, you can also use the interior-point and active-set algorithms of the fmincon solver. Specify the algorithm in the SearchOptions.Algorithm option. The fmincon algorithms might result in improved estimation results in the following scenarios:

  • Constrained minimization problems when bounds are imposed on the model parameters.

  • Model structures where the loss function is a nonlinear or nonsmooth function of the parameters.

  • Multiple-output model estimation. A determinant loss function is minimized by default for multiple-output model estimation. fmincon algorithms are able to minimize such loss functions directly. The other search methods such as 'lm' and 'gn' minimize the determinant loss function by alternately estimating the noise variance and reducing the loss value for a given noise variance value. Hence, the fmincon algorithms can offer better efficiency and accuracy for multiple-output model estimations.

'adam'

Adaptive moment estimation (Adam)

Adam is a first-order adaptive gradient solver. It supports mini-batch operation when data is segmented into multiple frames or batches. For more information, see Adaptive Moment Estimation (Deep Learning Toolbox).

'sgdm'

Stochastic gradient descent with momentum (SGDM)

SGDM is a first-order momentum-based gradient solver. It supports mini-batch operation when data is segmented into multiple frames or batches. For more information, see Stochastic Gradient Descent with Momentum (Deep Learning Toolbox).

'lbfgs'

Limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS)

L-BFGS is a quasi-Newton solver that approximates the inverse Hessian using a limited history of curvature pairs. For more information, see Limited-Memory BFGS (Deep Learning Toolbox).

Option set for the search algorithm, specified as a search option set with fields that depend on the value of SearchMethod.

SearchOptions Structure When SearchMethod Is Specified as 'gn', 'gna', 'lm', 'grad', or 'auto'

Field NameDescriptionDefault
Tolerance

Minimum percentage difference between the current value of the loss function and its expected improvement after the next iteration, specified as a positive scalar. When the percentage of expected improvement is less than Tolerance, the iterations stop. The estimate of the expected loss-function improvement at the next iteration is based on the Gauss-Newton vector computed for the current parameter value.

0.01
MaxIterations

Maximum number of iterations during loss-function minimization, specified as a positive integer. The iterations stop when MaxIterations is reached or another stopping criterion is satisfied, such as Tolerance.

Setting MaxIterations = 0 returns the result of the start-up procedure.

Use sys.Report.Termination.Iterations to get the actual number of iterations during an estimation, where sys is an idtf model.

20
Advanced

Advanced search settings, specified as a structure with the following fields.

Field NameDescriptionDefault
GnPinvConstant

Jacobian matrix singular value threshold, specified as a positive scalar. Singular values of the Jacobian matrix that are smaller than GnPinvConstant*max(size(J)*norm(J)*eps) are discarded when computing the search direction. Applicable when SearchMethod is 'gn'.

10000
InitialGnaTolerance

Initial value of gamma, specified as a positive scalar. Applicable when SearchMethod is 'gna'.

0.0001
LMStartValue

Starting value of search-direction length d in the Levenberg-Marquardt method, specified as a positive scalar. Applicable when SearchMethod is 'lm'.

0.001
LMStep

Size of the Levenberg-Marquardt step, specified as a positive integer. The next value of the search-direction length d in the Levenberg-Marquardt method is LMStep times the previous one. Applicable when SearchMethod is 'lm'.

2
MaxBisections

Maximum number of bisections used for line search along the search direction, specified as a positive integer.

25
MaxFunctionEvaluations

Maximum number of calls to the model file, specified as a positive integer. Iterations stop if the number of calls to the model file exceeds this value.

Inf
MinParameterChange

Smallest parameter update allowed per iteration, specified as a nonnegative scalar.

0
RelativeImprovement

Relative improvement threshold, specified as a nonnegative scalar. Iterations stop if the relative improvement of the criterion function is less than this value.

0
StepReduction

Step reduction factor, specified as a positive scalar that is greater than 1. The suggested parameter update is reduced by the factor StepReduction after each try. This reduction continues until MaxBisections tries are completed or a lower value of the criterion function is obtained.

StepReduction is not applicable for a SearchMethod of 'lm' (Levenberg-Marquardt method).

2

SearchOptions Structure When SearchMethod Is Specified as 'lsqnonlin'

Field NameDescriptionDefault
FunctionTolerance

Termination tolerance on the loss function that the software minimizes to determine the estimated parameter values, specified as a positive scalar.

The value of FunctionTolerance is the same as that of opt.SearchOptions.Advanced.TolFun.

1e-5
StepTolerance

Termination tolerance on the estimated parameter values, specified as a positive scalar.

The value of StepTolerance is the same as that of opt.SearchOptions.Advanced.TolX.

1e-6
MaxIterations

Maximum number of iterations during loss-function minimization, specified as a positive integer. The iterations stop when MaxIterations is reached or another stopping criterion is satisfied, such as FunctionTolerance.

The value of MaxIterations is the same as that of opt.SearchOptions.Advanced.MaxIter.

20

SearchOptions Structure When SearchMethod Is Specified as 'patternsearch'

Field NameDescriptionDefault
Algorithm

patternsearch optimization algorithm, specified as one of these values:

  • 'classic'

  • 'nups'

  • 'nups-gps'

  • 'nups-mads'

For algorithm details, see How Pattern Search Polling Works (Global Optimization Toolbox) and Nonuniform Pattern Search (NUPS) Algorithm (Global Optimization Toolbox).

For examples of algorithm effects, see Explore patternsearch Algorithms (Global Optimization Toolbox) and Explore patternsearch Algorithms in Optimize Live Editor Task (Global Optimization Toolbox).

'nups'
FunctionTolerance

Termination tolerance on the loss function that the software minimizes to determine the estimated parameter values, specified as a positive scalar.

1e-6
StepTolerance

Termination tolerance on the estimated parameter values, specified as a positive scalar.

1e-6
MaxIterations

Maximum number of iterations during loss function minimization, specified as a positive integer. The iterations stop when MaxIterations is reached or another stopping criterion is satisfied, such as FunctionTolerance.

'100*numberOfVariables', where numberOfVariables is the number of problem variables
UseParallel

Option to enable or disable parallel processing for improved performance, specified as one of these values:

  • "off" — Run in serial on the MATLAB® client.

  • "auto" — Use a parallel pool if one is open or if MATLAB can automatically create one. If a parallel pool is not available, run in serial on the MATLAB client.

  • "on" — Use a parallel pool if one is open or if MATLAB can automatically create one. If a parallel pool is not available, throw an error.

If you do not have a parallel pool open and automatic pool creation is enabled, MATLAB opens a pool using the default cluster profile. To use a parallel pool to run computations in MATLAB, you must have Parallel Computing Toolbox™.

Before R2026b: To run in parallel, set UseParallel to true.

"off"

SearchOptions Structure When SearchMethod Is Specified as 'fmincon'

Field NameDescriptionDefault
Algorithm

fmincon optimization algorithm, specified as one of the following:

  • 'sqp' — Sequential quadratic programming algorithm. The algorithm satisfies bounds at all iterations, and it can recover from NaN or Inf results. It is not a large-scale algorithm. For more information, see Sparsity in Optimization Algorithms (Optimization Toolbox).

  • 'trust-region-reflective' — Subspace trust-region method based on the interior-reflective Newton method. It is a large-scale algorithm.

  • 'interior-point' — Large-scale algorithm that requires Optimization Toolbox software. The algorithm satisfies bounds at all iterations, and it can recover from NaN or Inf results.

  • 'active-set' — Requires Optimization Toolbox software. The algorithm can take large steps, which adds speed. It is not a large-scale algorithm.

For more information about the algorithms, see Constrained Nonlinear Optimization Algorithms (Optimization Toolbox) and Choosing the Algorithm (Optimization Toolbox).

'sqp'
FunctionTolerance

Termination tolerance on the loss function that the software minimizes to determine the estimated parameter values, specified as a positive scalar.

1e-6
StepTolerance

Termination tolerance on the estimated parameter values, specified as a positive scalar.

1e-6
MaxIterations

Maximum number of iterations during loss function minimization, specified as a positive integer. The iterations stop when MaxIterations is reached or another stopping criterion is satisfied, such as FunctionTolerance.

100

SearchOptions Structure When SearchMethod Is Specified as 'adam'

Field NameDescriptionDefault
LearnRate

Learning rate, or the step size, used for training, specified as a positive scalar. If the learning rate is too small, then training can take a long time. If the learning rate is too large, then training can be fast but it might reach a suboptimal result, diverge, or oscillate. The learning rate is denoted by α in the Adaptive Moment Estimation (Deep Learning Toolbox) section.

If you specify LearnRateSchedule as "piecewise", then LearnRate is the learning rate before any scheduled drops.

0.001
GradientDecayFactor

Exponential decay rate of gradient moving average for the Adam solver, specified as a positive scalar less than 1. It controls the smoothing of the exponentially decaying average of past gradients. The gradient decay rate is denoted by β1 in the Adaptive Moment Estimation (Deep Learning Toolbox) section.

If the value of GradientDecayFactor is closer to 1, then the smoothing increases. If the value of GradientDecayFactor is closer to 0, then recent gradients have more impact on the training.

0.9
SquaredGradientDecayFactor

Exponential decay rate of squared gradient moving average for the Adam solver, specified as a positive scalar less than 1. It controls the smoothing of the exponentially decaying average of past squared gradients. The squared gradient decay rate is denoted by β2 in the Adaptive Moment Estimation (Deep Learning Toolbox) section.

Larger values of SquaredGradientDecayFactor adapt more slowly but provide a more stable variance estimate.

0.999
EpsilonSmall constant for numerical stability, specified as a positive scalar. To avoid division by zero when updating network parameters, the solver adds this constant to the denominator. Epsilon is denoted by ϵ in the Adaptive Moment Estimation (Deep Learning Toolbox) section.1e-8
MaxEpochsMaximum number of parameter updates or iterations to use for training, specified as a nonnegative integer. If you specify MaxEpochs as 0, the software disables iterations and only runs initialization or post-processing.200
MaxFunctionEvaluationsMaximum number of objective function evaluations, specified as a positive integer.intmax
LearnRateSchedule

Learning rate schedule type, specified as "none" or "piecewise".

  • "none" — No learning rate schedule. This schedule keeps the learning rate constant and equal to LearnRate.

  • "piecewise" — Piecewise learning rate schedule. This schedule multiplies the learning rate by LearnRateDropFactor every LearnRateDropPeriod number of iterations.

"none"
LearnRateDropFactor

Multiplicative factor for dropping the learning rate, specified as a positive scalar less than or equal to 1. You can specify this option only when you specify the LearnRateSchedule training option as "piecewise".

The software multiplies the learning rate with the factor specified by LearnRateDropFactor every time a certain number of iterations passes. Specify the number of iterations using the LearnRateDropPeriod training option.

0.1
LearnRateDropPeriod

Number of iterations in between learning rate drops, specified as a positive integer. You can specify this option only when you specify the LearnRateSchedule training option as "piecewise".

The software multiplies the learning rate with the drop factor every time the number of iterations specified by LearnRateDropPeriod passes. Specify the drop factor using the LearnRateDropFactor training option.

10
MinCostValue

Target objective value, specified as a nonnegative scalar. If the objective value at the current iteration is less than or equal to MinCostValue, the training stops.

0
ModelSelection

Iteration used to return the model parameters, specified as "best" or "last".

If you specify ModelSelection as "best", the software returns the parameters corresponding to the iteration with the lowest objective value. If you specify ModelSelection as "last", the software returns the parameters corresponding to the final iteration.

"best"
AdvancedStructure used to specify advanced search options consisting of these fields:

WeightDecay — Strength of L2 regularization applied to learnable parameters, specified as a nonnegative scalar.

  • If WeightDecay>0 and UseDecoupledWeightDecay=1, weight decay is applied in a decoupled manner before update computations.

  • If WeightDecay>0 and UseDecoupledWeightDecay=0, classical L2 regularization is applied by augmenting gradients before update computations.

To disable this option, specify WeightDecay as 0.

0

UseDecoupledWeightDecay — Flag to control whether weight decay is applied directly to parameters or is folded into the gradient as L2 penalty, specified as a logical scalar.

  • If UseDecoupledWeightDecay=1, decoupled weight decay is applied directly.

  • If UseDecoupledWeightDecay=0, L2 penalty term is added to the gradient.

Decoupling prevents regularization strength from being implicitly modulated by momentum dynamics.

1

ClipGradNorm — Maximum allowed L2 norm of the flattened gradient vector, specified as a nonnegative scalar. If the gradient norm exceeds this value, the gradient is rescaled. This rescaling helps prevent unstable updates when gradients spike, such as in recurrent models or stiff dynamical systems.

To disable this option, specify ClipGradNorm as 0.

0
NormEpsilon — Small constant used to avoid division by zero and underflow when computing safe norms and normalized quantities, especially when ClipGradNorm is enabled, specified as a positive scalar.1e-12

SearchOptions Structure When SearchMethod Is Specified as 'sgdm'

Field NameDescriptionDefault
LearnRate

Learning rate, or the step size, used for training, specified as a positive scalar. If the learning rate is too small, then training can take a long time. If the learning rate is too large, then training can be fast but it might reach a suboptimal result, diverge, or oscillate. The learning rate is denoted by α in the Stochastic Gradient Descent with Momentum (Deep Learning Toolbox) section.

If you specify LearnRateSchedule as "piecewise", then LearnRate is the learning rate before any scheduled drops.

0.01
Momentum

Momentum coefficient, specified as a positive scalar less than or equal to 1. This coefficient controls the contribution of the previous gradient step to the current iteration. The momentum coefficient is denoted by γ in the Stochastic Gradient Descent with Momentum (Deep Learning Toolbox) section.

If the value of Momentum is closer to 1, then the smoothing increases. If the value of Momentum is closer to 0, then the solver behaves closer to the stochastic gradient descent algorithm.

0.95
MaxEpochsMaximum number of parameter updates or iterations to use for training, specified as a nonnegative integer. If you specify MaxEpochs as 0, the software disables iterations and only runs initialization or post-processing.200
MaxFunctionEvaluationsMaximum number of objective function evaluations, specified as a positive integer.intmax
LearnRateSchedule

Learning rate schedule type, specified as "none" or "piecewise".

  • "none" — No learning rate schedule. This schedule keeps the learning rate constant and equal to LearnRate.

  • "piecewise" — Piecewise learning rate schedule. This schedule multiplies the learning rate by LearnRateDropFactor every LearnRateDropPeriod number of iterations.

"none"
LearnRateDropFactor

Multiplicative factor for dropping the learning rate, specified as a positive scalar less than or equal to 1. You can specify this option only when you specify the LearnRateSchedule training option as "piecewise".

The software multiplies the learning rate with the factor specified by LearnRateDropFactor every time a certain number of iterations passes. Specify the number of iterations using the LearnRateDropPeriod training option.

0.1
LearnRateDropPeriod

Number of iterations in between learning rate drops, specified as a positive integer. You can specify this option only when you specify the LearnRateSchedule training option as "piecewise".

The software multiplies the learning rate with the drop factor every time the number of iterations specified by LearnRateDropPeriod passes. Specify the drop factor using the LearnRateDropFactor training option.

10
MinCostValue

Target objective value, specified as a nonnegative scalar. If the objective value at the current iteration is less than or equal to MinCostValue, the training stops.

0
ModelSelection

Iteration used to return the model parameters, specified as "best" or "last".

If you specify ModelSelection as "best", the software returns the parameters corresponding to the iteration with the lowest objective value. If you specify ModelSelection as "last", the software returns the parameters corresponding to the final iteration.

"best"
AdvancedStructure used to specify advanced search options consisting of these fields:

WeightDecay — Strength of L2 regularization applied to learnable parameters, specified as a nonnegative scalar.

  • If WeightDecay>0 and UseDecoupledWeightDecay=1, weight decay is applied in a decoupled manner before update computations.

  • If WeightDecay>0 and UseDecoupledWeightDecay=0, classical L2 regularization is applied by augmenting gradients before update computations.

To disable this option, specify WeightDecay as 0.

0

UseDecoupledWeightDecay — Flag to control whether weight decay is applied directly to parameters or is folded into the gradient as L2 penalty, specified as a logical scalar.

  • If UseDecoupledWeightDecay=1, decoupled weight decay is applied directly.

  • If UseDecoupledWeightDecay=0, L2 penalty term is added to the gradient.

Decoupling prevents regularization strength from being implicitly modulated by momentum dynamics.

1

ClipGradNorm — Maximum allowed L2 norm of the flattened gradient vector, specified as a nonnegative scalar. If the gradient norm exceeds this value, the gradient is rescaled. This rescaling helps prevent unstable updates when gradients spike, such as in recurrent models or stiff dynamical systems.

To disable this option, specify ClipGradNorm as 0.

0
NormEpsilon — Small constant used to avoid division by zero and underflow when computing safe norms and normalized quantities, especially when ClipGradNorm is enabled, specified as a positive scalar.1e-12

SearchOptions Structure When SearchMethod Is Specified as 'lbfgs'

Field NameDescriptionDefault
MaxIterations

Maximum number of quasi-Newton iterations to use for training, specified as a nonnegative integer. Each iteration forms a search direction using the stored curvature pairs and then performs a line search.

If you specify MaxIterations as 0, the software disables iterations and only runs initialization or post-processing.

200
MaxFunctionEvaluationsMaximum number of objective function evaluations, including evaluations performed by line search, specified as a positive integer.intmax
HistorySize

Number of curvature pairs or state updates to store, specified as a positive integer.

The L-BFGS algorithm uses a history of gradient calculations to approximate the Hessian matrix recursively. Larger values of HistorySize can improve the Hessian approximation but will increase memory usage and cost per iteration. For more information, see the Limited-Memory BFGS (Deep Learning Toolbox) section.

10
GradientTolerance

Stopping tolerance on the relative gradient, specified as a positive scalar.

The software stops training when the relative gradient is less than or equal to GradientTolerance.

1e-6
StepTolerance

Stopping tolerance on the step size, specified as a positive scalar. StepTolerance specifies the minimum allowable change in the parameters between successive iterations.

The software stops training when the step that the algorithm takes is less than or equal to StepTolerance.

1e-12
FunctionTolerance

Stopping tolerance on the improvement in the objective value, specified as a positive scalar. FunctionTolerance specifies the minimum required decrease in the objective function value between iterations.

The software stops training when the objective value improvement is less than or equal to FunctionTolerance.

1e-12
LineSearchMethod

Method to find a suitable step size, specified as one of these values:

  • "strong-wolfe" — Search for a step size that satisfies the strong Wolfe conditions (sufficient decrease and strong curvature). This method maintains a positive definite approximation of the inverse Hessian matrix.

  • "weak-wolfe" — Search for a step size that satisfies the weak Wolfe conditions (sufficient decrease and curvature). This method maintains a positive definite approximation of the inverse Hessian matrix. It can accept longer steps.

  • "armijo" — Search for a learning rate that satisfies sufficient decrease conditions only. This method does not maintain a positive definite approximation of the inverse Hessian matrix. It is often more tolerant of noisy gradients but can accept shorter steps.

"strong-wolfe"
MaxNumLineSearchIterationsMaximum number of line search trials per iteration to determine the step size, specified as a positive integer.40
InitialStepSizeStep size for the starting line search trial, specified as a positive scalar.1.0
AdvancedStructure used to specify advanced search options consisting of these fields:
MinStepSize — Smallest step size permitted by line search, specified as a positive scalar. If the step size for a trial goes below this value, the line search fails and the solver can stop or fall back depending on the implementation.1e-16
MaxStepSize — Largest step size permitted by line search, specified as a positive scalar. This upper bound for the trial step size prevents excessively large moves that can cause numerical overflow or objective evaluation failures.1e+16

GradientClipNorm — Maximum allowed L2 norm of the gradient vector used by the quasi-Newton update and line search, specified as a nonnegative scalar. If the gradient norm exceeds this value, the gradient is rescaled. This rescaling helps improve robustness on problems with occasional gradient spikes or poor scaling.

To disable this option, specify GradientClipNorm as 0.

0
WolfeC1 — Armijo condition (sufficient decrease) constant for Wolfe line search, specified as a positive scalar less than 1. Smaller values of WolfeC1 make sufficient decrease easier to satisfy.1e-4
WolfeC2 — Curvature condition constant for Wolfe line search, specified as positive scalar less than 1. Larger values of WolfeC2 make the curvature condition easier to satisfy whereas smaller values enforce a stronger curvature requirement.0.9
ZoomMaxIterations — Maximum number of iterations allowed in the "zoom" procedure of Wolfe line search, specified as a positive integer.40
BacktrackingFactor — Step size shrink factor during backtracking used to reduce trial step sizes when conditions are not satisfied, specified as a positive scalar less than 1. Values closer to 0 shrink the step size more aggressively while values closer to 1 shrink the step size more conservatively.0.5
CurvatureThreshold — Number to control whether a new curvature pair is accepted into the limited-memory history, specified as a positive scalar. Specifying CurvatureThreshold prevents storing nearly singular or noisy curvature information that can destabilize the inverse-Hessian approximation.1e-10

PowellDamping — Number to control the amount of Powell damping applied when the curvature condition is weak, specified as a nonnegative number less than 1. Damping enforces positive curvature and a positive-definite inverse-Hessian approximation.

To disable this option, specify PowellDamping as 0.

0
UseInitialScaling — Flag to control whether the initial inverse-Hessian is scaled each iteration using curvature information, specified as a logical scalar. This scaling often improves practical performance.1

Additional advanced options, specified as a structure with the following fields:

  • ErrorThreshold — Specifies when to adjust the weight of large errors from quadratic to linear.

    Errors larger than ErrorThreshold times the estimated standard deviation have a linear weight in the loss function. The standard deviation is estimated robustly as the median of the absolute deviations from the median of the prediction errors and divided by 0.7. For more information on robust norm choices, see section 15.2 of [2].

    ErrorThreshold = 0 disables robustification and leads to a purely quadratic loss function. When estimating with frequency-domain data, the software sets ErrorThreshold to zero. For time-domain data that contains outliers, try setting ErrorThreshold to 1.6.

    Default: 0

  • MaxSize — Specifies the maximum number of elements in a segment when input-output data is split into segments.

    MaxSize must be a positive integer.

    Default: 250000

  • StabilityThreshold — Specifies thresholds for stability tests.

    StabilityThreshold is a structure with the following fields:

    • s — Specifies the location of the right-most pole to test the stability of continuous-time models. A model is considered stable when its right-most pole is to the left of s.

      Default: 0

    • z — Specifies the maximum distance of all poles from the origin to test stability of discrete-time models. A model is considered stable if all poles are within the distance z from the origin.

      Default: 1+sqrt(eps)

  • AutoInitThreshold — Specifies when to automatically estimate the initial condition.

    The initial condition is estimated when

    yp,zymeasyp,eymeas>AutoInitThreshold

    • ymeas is the measured output.

    • yp,z is the predicted output of a model estimated using zero initial conditions.

    • yp,e is the predicted output of a model estimated using estimated initial conditions.

    Applicable when InitialCondition is 'auto'.

    Default: 1.05

Examples

collapse all

opt = armaxOptions;

Create an option set for armax to use the 'simulation' Focus and to set the Display to 'on'.

opt = armaxOptions('Focus','simulation','Display','on');

Alternatively, use dot notation to set the values of opt.

opt = armaxOptions;
opt.Focus = 'simulation';
opt.Display = 'on';

References

[1] Wills, Adrian, B. Ninness, and S. Gibson. “On Gradient-Based Search for Multivariable System Estimates”. Proceedings of the 16th IFAC World Congress, Prague, Czech Republic, July 3–8, 2005. Oxford, UK: Elsevier Ltd., 2005.

[2] Ljung, L. System Identification: Theory for the User. Upper Saddle River, NJ: Prentice-Hall PTR, 1999.

Extended Capabilities

expand all

Version History

Introduced in R2012a

expand all