SCORISTA Hotline +7 (495) 268-09-63
Model errors occur in credit scoring: situations where the model perceives a potentially reliable borrower as uncreditworthy, or conversely, evaluates an uncreditworthy borrower as "good" (Type I and Type II errors). Such errors can lead to significant financial losses associated with missed profit opportunities.
There are many reasons why a scoring model can be wrong. Key factors include model overfitting and insufficient training sample sizes. Furthermore, scoring models are static, whereas the constantly evolving borrower profile and changing lending conditions lead to performance instability. It is also worth noting that building a scoring model is a labor-intensive and resource-heavy process requiring substantial computing power and time. At times, there is a need to rapidly adapt to new realities while maintaining the quality of key performance indicators. To address this, Scorista implemented a method known as "String Segmentation."
String segmentation is a decision-making approach involving a multidimensional assessment of the borrower, enabling the precise minimization of scoring model errors. The essence of the method lies in the comprehensive integration of several mutually independent models. This superposition of models enhances decision-making reliability and positively impacts key metrics, such as zero-default rates and FPD (First Payment Default).
Let us examine the sequence of steps required to execute the string segmentation procedure:
Another key advantage of this approach is the ease with which the approval decision can be described. For the sake of simplicity and clarity, let us consider an example involving two models. Suppose we have a scoring model (Model 1) that has begun to make errors when assessing borrower creditworthiness. Incorporating a second model (Model 2) into the decision process significantly improved discriminatory power (Fig. 1). It became evident that Model 1 was producing both Type I and Type II errors.

Fig. 1. FPD15 metric plotted against the scores from two uncorrelated scoring models.
Of course, in a two-model scenario, one could painstakingly manually define a multitude of conditions determining which score combinations from the two models would result in approval or rejection. But what if we wanted to increase the decision's granularity further and use three or more models? In that case, the approval decision formula would involve countless rules and exceptions. When using string-based segmentation, the problem of describing every exception vanishes, and the approval decision itself can be expressed in just a couple of lines.
Fig. 2 illustrates how the method works. Each of the two models was divided into 19 bins. The bin identifiers were combined into a single string variable (*sv*).

Fig. 2. The string variable and its values resulting from the combination of two models.
Next, the resulting string variables are grouped into *n* categories based on the target variable, and specific approval groups are selected that meet the required quantitative and qualitative criteria. The identifiers of these resulting groups constitute the approval decision itself.
The two-model example shown here was chosen solely for the sake of simplicity and visual clarity. In practice, at Skorista, at least three different uncorrelated models are used for string segmentation.
We compare the performance of a classic model aggregation method (stacking) with the string segmentation method using real-world data.
A microfinance company issues short-term payday loans (PDL) with a 16-day term. A decision-making model—Model A (Fig. 3)—is used in the approval process; a score above 500 satisfies one of the key requirements: approving 22% of the borrower flow.
Fig. 3. A metamodel obtained using a traditional aggregation method.
This metamodel was derived from four models (Fig. 4) built using logistic regression. Each of these models demonstrates high predictive power. The metamodel was aggregated using a traditional machine learning technique: stacking.

Fig. 4. Models built using logistic regression.
To ensure a fair comparison, we now perform the String Segmentation procedure using the same four models (Fig. 4).
| Bin size | Number of models in the string variable | Group size | Number of groups | Stabilization parameter |
|---|---|---|---|---|
| 20% | 4 | 5% | 5 | 80% |
Table 1. Optimal parameters for string segmentation.

Fig. 5. String segmentation grouping results. Approval of groups 1 and 2.
We now compare the results obtained using the traditional model aggregation method and the String Segmentation method over the test period (Fig. 6).
Fig. 6. FPD15 and zero-default metrics on the test sample, where 0 = rejection and 1 = approval.
Figure 6 demonstrates the advantage of String Segmentation over the traditional model aggregation method. The use of String Segmentation significantly improved key metrics while maintaining the same approval rate.
The example above illustrates the use of String Segmentation as a standalone method for making approval decisions. This method also yielded strong results when combined with conventional model aggregation techniques—serving as an additional borrower assessment that enhances the quality of the primary decision.
In conclusion, String Segmentation is a method that enables high-quality approval decisions.
Its key advantages include:
Independent* – in the context of this article, model independence refers to minimal correlation between the models (based on Cramer's V).