Data Science
•
•
5 min read

How GBMs Work

A practical explanation of gradient boosting machines, pseudo-residuals, loss functions, feature importance, missing values, multicollinearity, and quantile GBM.

GBM Machine Learning Modeling

As the name suggests, GBM can be understood through two concepts:

Gradient Boosting = Gradient Descent + Boosting

It builds an ensemble of regression trees sequentially, where each new tree tries to correct the errors of the current ensemble rather than starting from scratch.

Each tree is trained to approximate the negative gradient of the loss function with respect to the current predictions, and these negative gradients are often called pseudo-residuals.

The negative gradient tells us how each current prediction should change to reduce the loss. The weak CART tree learns a function that approximates those desired changes.

Current ensemble
        │
        ▼
Compute negative gradient for every row
        │
        ▼
These become regression targets
        │
        ▼
Train an ordinary CART regression tree
        │
        ├── Try Age < 40
        ├── Compute RSS reduction
        ├── Try Income < 60k
        ├── Compute RSS reduction
        └── Choose the split with the largest RSS reduction
        │
        ▼
Add the tree to the ensemble
        │
        ▼
Repeat

The Core Idea

In linear regression, gradient descent updates coefficients. The gradient vector is tied to the columns of the design matrix because the parameters are the coefficients.

In GBM, the model is updated through functions. The gradient vector has size n, the number of training observations, because the gradient is taken with respect to the current prediction for each row. The tree converts these row-level gradient values into a function that can update predictions:

F_m(x) = F_{m-1}(x) + nu * h_m(x)

where:

  • F_{m-1}(x) is the current ensemble
  • h_m(x) is the new weak learner
  • nu is the learning rate

Now that the weak learner is the best approximation of the gradient of the loss function, we update the previous ensemble by adding the weak learner, multiplied by the step size. This is mimicking gradient descent, but instead of updating coefficients directly, GBM adds weak learners iteratively.

For squared-error loss, the pseudo-residuals are the ordinary residuals. For other loss functions, they are the corresponding negative gradients.

This is also one of the key differences from Random Forest: Random Forest builds independent trees, while GBM builds each tree to optimize a single global objective.

How GBM Works Step by Step

  1. Start with an initial prediction.
  2. Make predictions using the current model.
  3. Compute pseudo-residuals.
  4. Train a weak learner on the pseudo-residuals.
  5. Scale the new tree by the learning rate.
  6. Repeat.
  7. Make the final prediction.

The final model is the sum of all corrections added to the initial prediction:

F_M(x) = F_0(x) + nu * sum_{m=1}^{M} h_m(x)

Loss Function

The boosting algorithm chooses the loss function based on the prediction task.

  • Regression: squared error, absolute error, Huber loss
  • Binary classification: logistic loss or cross-entropy
  • Count data: Poisson deviance
  • Positive continuous data: Gamma deviance or Tweedie deviance

The chosen loss determines the pseudo-residuals by computing the negative gradient of the loss with respect to the current predictions.

Important note: the individual regression tree does not directly optimize the original loss function. The CART tree is fit to the pseudo-residuals by greedily choosing splits that minimize Residual Sum of Squares (RSS).

So GBM is nonparametric in the sense that it does not assume a specific functional form for the prediction function. But it still requires a loss function to define what a good prediction means. That loss is often chosen based on the response variable, such as Gaussian, Bernoulli, Poisson, Gamma, or Tweedie.

Number of Trees and Early Stopping

Tree-based boosting is iterative, so model complexity increases as trees are added and there is no natural stopping point.

That is why a validation set is important during training, since we need to evaluate out-of-sample performance after each iteration.

Early stopping monitors a validation metric and stops training once performance no longer improves for a specified number of rounds. The final model is usually selected from the iteration with the best validation score.

Feature Importance

GBMs can calculate feature importance in a few different ways.

  • Split Gain

    • Measures the total reduction in RSS of the pseudo-residuals whenever a feature is used for a split.
    • Features that produce larger reductions in RSS receive higher importance.
    • The gain is accumulated over every tree in the ensemble.
  • Split Count

    • Counts how many times a feature is selected for splitting across all trees.
    • More splits imply greater importance.
  • Coverage

    • Measures how many training observations are affected by splits using a feature. Features used near the top of trees usually have higher coverage.

Missing Values

FrameworkHow it handles missing values
XGBoostAutomatically learns the best default direction for missing values at each split based on the objective or split gain.
LightGBMHandles missing values natively and determines the best direction during split optimization.
CatBoostHandles numerical missing values natively and can consider splits that separate missing values from non-missing values.
Older sklearn GradientBoostingDoes not natively handle missing values; usually requires imputation or removal.

Multicollinearity

GBMs are generally robust to multicollinearity for prediction, although correlated features can make feature importance and interpretation unstable.

Unlike GLMs, GBMs do not estimate feature coefficients through matrix inversion. Trees greedily select splits based on which feature provides the greatest improvement in the objective. So highly correlated predictors do not create unstable coefficient estimates, because there are no feature coefficients to estimate.

If X1 and X2 contain similar information, the GBM can use either and usually achieve similar predictive performance; the issue is attribution.

Correlated features compete for splits. If feature A is selected first and captures most of the signal shared with feature B, then feature B may have little additional gain left:

A selected -> shared signal captured -> B has little incremental gain -> B selected less often

Because tree building is greedy, small differences in gain, data, feature subsampling, or implementation details can cause one correlated feature to dominate while another is rarely used.

That means zero feature importance does not necessarily mean zero predictive information, because another correlated feature may have captured the same signal.

Permutation importance and SHAP can provide additional insight, but correlation also affects their interpretation, so they help without completely solving the attribution problem.

AspectGLMGBM
Effect of multicollinearityCoefficients and standard errors can become unstable.Correlated features compete for splits.
PredictionOften remains good if train/test relationships are similar.Generally remains good.
Main concernCoefficient inference and interpretation.Feature importance and interpretation.
WhyIt is difficult to uniquely attribute shared signal to individual coefficients.It is difficult to uniquely attribute shared signal to individual features.
Potential prediction issueCan become unstable when extrapolating to new combinations of correlated features.Usually less sensitive, although distribution shift can still matter.

Interactions

Decision trees naturally capture interactions. When a tree splits on feature A, it partitions the data into groups where the relationship between other features and the target may be different. If the tree then splits on feature B within one of those partitions, it has effectively captured an interaction between A and B.

This is why tree depth matters: deeper trees can capture more complex interactions, but they also increase the risk of overfitting.

Regularization

GBMs can overfit if the model is allowed to keep adding flexible trees without control, so regularization is what keeps the model useful.

Common regularization tools include:

  • Row subsampling: use a random fraction of training observations for each tree, such as subsample = 0.8.
  • Tree complexity limits: restrict depth, leaf count, or minimum observations per leaf.
  • Learning rate: shrink each tree’s contribution.
  • Early stopping: stop adding trees when validation performance stops improving.
  • L1 and L2 penalties: XGBoost can apply explicit penalties to leaf weights.

Classical GBM does not usually include explicit L1/L2 regularization, while Random Forest has a different kind of overfitting resistance through bagging and feature randomness.

Quantile Gradient Boosting

Quantile Gradient Boosting predicts a conditional quantile of the target rather than the conditional mean, which makes it useful when I care about a range of possible outcomes instead of a single expected value.

Standard GBM with squared-error loss estimates:

E[Y | X]

Quantile GBM estimates values like:

Q_tau(Y | X)

where:

  • tau = 0.5 estimates the median
  • tau = 0.1 estimates the 10th percentile
  • tau = 0.9 estimates the 90th percentile

The quantile loss function is:

L_tau(y, y_hat) =
    tau * (y - y_hat)          if y >= y_hat
    (1 - tau) * (y_hat - y)    if y < y_hat

The loss penalizes underprediction and overprediction asymmetrically. For example, when tau = 0.9, underprediction (y^ < y) is penalized by 0.9, while overprediction (y^ > y) is penalized by 0.1; as a result, the model favors predictions near the upper tail of the conditional distribution because underestimating large values is more costly.

When tau = 0.1, overprediction is more costly, so the model produces more conservative lower-tail predictions.

The boosting procedure itself is unchanged because trees are still added sequentially; the only difference is that errors are measured using quantile loss instead of squared error.

Quantile GBMs are useful when uncertainty matters. By training models at multiple quantiles, such as the 10th, 50th, and 90th percentiles, we can construct risk-aware prediction bands rather than relying on a single point estimate.

lower band      median prediction      upper band
Q_0.10(Y | X)   Q_0.50(Y | X)          Q_0.90(Y | X)

      |--------------- prediction band ---------------|