偏差和方差
Suppose there is an unknown target function or “true function” f(x) that maps input vector X to output Y. For instance, f(X) could be the function that takes features of the house (#no of the bedroom, distance to the nearest hospital, etc.)and maps it to its corresponding price Y.
小号uppose则存在输入矢量X映射到输出Y.一个未知的目标函数或“真函数” f(x)的。 例如, f(X)可以是获取房屋特征(卧室编号,到最近医院的距离等)并将其映射到其对应价格Y的函数。
For any given input X there might not exist a unique label Y. we could imagine two houses with identical features but with a different price. This randomness could be due to various factors and that affects Y but we haven't taken into account with our function f. The relationship between the seller and owner can be one of such factors. We represent such randomness with error terms ϵ.
对于任何给定的输入X ,可能不存在唯一标签Y。 我们可以想象两座功能相同但价格不同的房屋。 这种随机性可能是由于各种因素造成的,并且影响了Y,但是我们并未将函数f考虑在内。 卖方和所有者之间的关系可以是这样的因素之一。 我们用误差项ϵ表示这种随机性。
Here the choice of function f could be linear, in which case we may estimate f via a linear regression model. It may be non-linear too, in which case we may estimate f with a Support Vector Machine.
在这里,函数f的选择可以是线性的,在这种情况下,我们可以通过线性回归模型来估计f 。 它也可能是非线性的,在这种情况下,我们可以使用支持向量机来估计f 。
Let's consider a hypothetical world where we know the true relationship f(X) between X and Y. Note that in reality, we will not ever know the underlying f, which is why we are estimating it in the first place. Suppose we observe some training dataset(say Dataset 1) from that relationship and use that to approximate true relationship. Let’s say it f1(x). However, we could repeat the whole model building process more than once: each time we gather new data (say Dataset 2)and run a new analysis and estimate a new approximation(say f2(X)) as shown in the below figure.
让我们考虑一个假设世界,在该世界中我们知道X和Y之间的真实关系f(X) 。 请注意,实际上,我们永远不会知道基础f ,这就是为什么我们首先对其进行估计的原因。 假设我们从该关系中观察到一些训练数据集(例如数据集1),并使用它来近似真实关系。 让我们 说 它 f1(x)。 但是,我们可以重复整个模型构建过程不止一次:每次我们收集新数据(例如数据集2)并运行新分析并估计新的近似值(例如f2(X)),如下图所示。
fig1: Illustration of approximated function based on different possible datasets from the same data generator 图1:基于来自同一数据生成器的不同可能数据集的近似函数的图示The following plot shows different linear regression models, each fit to a different training set where the green curve is the true function and red is the approximated function based on the individual datasets.
下图显示了不同的线性回归模型,每个模型都适合于不同的训练集,其中绿色曲线是真实函数,红色曲线是基于各个数据集的近似函数。
fig2: Illustration of linear model fit on the different possible datasets 图2:适用于不同可能数据集的线性模型拟合图Due to randomness in the underlying data sets, the resulting approximated linear function would be different too.
由于基础数据集中的随机性,所得的近似线性函数也将有所不同。
We can now try to fit a few different models other than linear to these different datasets.Let’s try to fit polynomial models on Datasets 1 as shown in the following figure:
我们现在可以尝试将线性模型以外的其他一些模型拟合到这些不同的数据集,让我们尝试将数据集1上的多项式模型拟合如下图所示:
Figure 3: Illustration of polynomial model fit on dataset 1 图3:拟合数据集1的多项式模型的图示In figure 3, the model, given by the blue curve, is a polynomial model with degree m=3. The second model, given by the red curve is a higher degree polynomial with degree m=10. Between each of the models, there is variation in the flexibility, that is, the degrees of freedom (DoF). The most flexible model is the polynomial of order m=10. It can be clearly seen, as the flexibility of the model increases it can fit better to a data point and the closest fit is the one with polynomial order m = 10. By allowing the model to be extremely flexible we are letting it fit to “patterns” in the training data.
在图3中,由蓝色曲线给出的模型是次数为m = 3的多项式模型。 红色曲线给出的第二个模型是次数为m = 10的高次多项式。 在每个模型之间,灵活性(即自由度(DoF))存在差异。 最灵活的模型是阶数为m = 10的多项式。 可以清楚地看到,随着模型灵活性的提高,它可以更好地拟合数据点,而最接近的拟合则是多项式阶数m = 10的拟合。通过允许模型具有极高的灵活性,我们使其适合于“训练数据中的“模式”。
However, if we try to fit the same model with m=10 that was trained on Dataset 1 on different datasets say Dataset 2 as shown in the below figure it will not perform well. In can be clearly seen in the following curve the model with polynomial order m = 3 fits better as compared to the model with m = 10. It is because sometimes these “patterns” are only random artifacts of the training data(Datasets 1) and are not an underlying property of the true function. Such problems in machine learning are called overfitting.
但是,如果我们尝试使用在不同数据集的数据集1上训练过的m = 10的相同模型,则说数据集2,如下图所示,它将无法很好地执行。 在下面的曲线中可以清楚地看到,与m = 10的模型相比,多项式为m = 3的模型更适合。这是因为有时这些“模式”只是训练数据(数据集1)和不是true函数的基础属性。 机器学习中的此类问题称为过拟合。
Figure 4: Illustration of model fit on unseen data (dataset 2) 图4:对看不见的数据进行模型拟合的图示(数据集2)In general, the approximated function would not be the perfect fit mainly due to the two sources of error.
通常,主要由于两个误差源,近似函数将不是最佳拟合。
Bias, which is caused by the choice of the model. In the above plot, it’s impossible for any linear function to exactly match the curve we’re looking for. 偏差是由模型的选择引起的。 在上面的图中,任何线性函数都不可能完全匹配我们要寻找的曲线。 The variance that comes from randomness inherent in the training set. 来自训练集中固有的随机性的方差。It is defined as the difference between the expected value of the parameter (or approximated function) and the population parameter (or true function) that we want to estimate. It is a way to measure how far on average we are from the ground truth.
它定义为参数(或近似函数)的期望值与我们要估计的总体参数(或真函数)之间的差。 这是一种衡量我们离基本事实平均距离的方法。
It is an error due to the classifier being “biased” to the particular solution example linear classifier. In other words, it is an error due to the choice of model and you will get this error even with infinite training data until and unless we change the choice of classifier.
由于分类器“偏向”特定解决方案示例线性分类器,因此出现错误。 换句话说,由于模型的选择,这是一个错误,即使在无数训练数据的情况下,您也会遇到此错误,除非并且除非我们更改分类器的选择。
In figure 2, we have tried to fit the linear model to our data points. Since the flexibility of the linear model is low we are far on average from ground truth function. If we try to fit the same data points using higher degree polynomial due to its more flexible in nature it can fit the data points quite well.
在图2中,我们尝试将线性模型拟合到我们的数据点。 由于线性模型的灵活性较低,因此我们离地面真实函数的平均距离还很远。 如果我们尝试使用高阶多项式拟合相同的数据点,因为它本质上更灵活,则可以很好地拟合数据点。
When we talk about variance we talk over many datasets. The error due to variance is taken as the variability of a model prediction between different datasets It is a way to measure how much our present model is overspecialized to a particular training set(overfitting).
当我们谈论方差时,我们会讨论许多数据集。 由于方差引起的误差被视为不同数据集之间模型预测的差异性。这是一种方法,用于衡量我们的当前模型对特定训练集的过度专业化程度(过度拟合)。
In figure 3, the polynomial model with order m = 10 is overspecialized to Dataset 1 and performs poorly on Dataset 2.So it is a model with high variance.
在图3中,阶数为m = 10的多项式模型专门针对数据集1,而在数据集2上的性能较差,因此它是具有高方差的模型。
We can graphically visualize the concept of bias and variance using the bulls-eye diagram. Imagine the center of the target is the value that perfectly predicts the model parameter. As we move away from the target our prediction gets worse and worse.
我们可以使用靶心图以图形方式形象化偏见和方差的概念。 假设目标的中心是完美预测模型参数的值。 随着我们远离目标,我们的预测越来越糟。
source) 来源)Imagine you are holding a gun and shooting a bunch of bullets on the target. You can think of each individual bullet as an estimator (colored as blue) build on different datasets such that each hit represents an individual estimation of the model parameters given their own datasets. The center of the target marked as red in the figure is the true population parameter that we are trying to estimate.
想象一下,您拿着枪在目标上射击了一堆子弹。 您可以将每个项目符号看作是建立在不同数据集上的估算器(颜色为蓝色),以便每个命中代表给定自己数据集的模型参数的单个估算。 图中标记为红色的目标的中心是我们试图估计的真实人口参数。
When we have a low bias, by the definition we are close to the ground truth. In this case, we are close to the center. Similarly, when we have low variance our estimator across different datasets is not going to be too spread out.
当我们的偏见低时,根据定义,我们接近基本事实。 在这种情况下,我们靠近中心。 同样,当方差低时,我们在不同数据集上的估计量也不会过于分散。
At its root, dealing with bias and variance is really about dealing with over- and under-fitting. When we try to fit the simple model it may not flexible enough in order to capture true relationships between input and output. In that case, we are oversimplifying the solution. Such a problem is called underfitting.
从根本上讲,处理偏差和方差实际上是处理过度拟合和拟合不足。 当我们尝试拟合简单模型时,它可能不够灵活,无法捕获输入和输出之间的真实关系。 在这种情况下,我们将简化解决方案。 这样的问题称为欠拟合。
However, when we increase the complexity of the model it may try to fit the noise inherent in that particular training datasets and perform worse on unseen data points. In that case, Variance becomes our primary concern. Such a problem is called underfitting.
但是,当我们增加模型的复杂度时,它可能会尝试拟合该特定训练数据集中固有的噪声,并对看不见的数据点表现更差。 在这种情况下,差异成为我们的首要关注。 这样的问题称为欠拟合。
The expected test MSE for a particular model estimate with the input vector X = x₀ where the expectation is taken across many training sets is given by:
对于具有输入向量X = x a的特定模型估计的期望测试MSE ,其中期望是跨许多训练集得出的:
However, we can expand the expectation on the right-hand side into three terms:
但是,我们可以将右侧的期望扩展为三个术语:
By the definition of bias and variance it can be expressed as:
根据偏差和方差的定义,可以表示为:
The third term, irreducible error, is the noise term in the true relationship that cannot fundamentally be reduced by any model. It is the minimum lower bound for the test MSE.
第三项,不可减少的误差,是真实关系中的噪声项,任何模型都无法从根本上减少噪声项。 它是测试MSE的最小下限。
Given the true model and infinite data to calibrate it, we should be able to reduce both the bias and variance terms to 0. However, in a world with imperfect models and finite data, there is a tradeoff between minimizing the bias and minimizing the variance.
给定真实的模型和用于校准的无限数据,我们应该能够将偏差项和方差项都减少到0。但是,在模型和有限数据不完善的世界中,需要在最小化偏差和最小化方差之间进行权衡。
Generally, as the complexity increases bias will drop quickly(faster than variance can increase)so we can see a drop in test error. However, when we further increase the complexity, there is less reduction in bias and a rapid increase in variance. Our ultimate goal is to find the model with the minimum expected test MSE.So we must choose a model with low variance as well as low bias.
通常,随着复杂度的增加,偏差将Swift下降(快于方差的增加),因此我们可以看到测试误差的下降。 但是,当我们进一步增加复杂性时,偏差的减少会减少,方差会Swift增加。 我们的最终目标是找到具有最小预期测试MSE的模型,因此我们必须选择方差低且偏差小的模型。
source) 来源)翻译自: https://medium.com/@regmi.sobit/bias-and-variance-clearly-explained-6d5efcdf076
偏差和方差
