根据nodes复习知识点 chatgpt 大哥挖掘面试问题Supervised learningLinear regression公开资料显示Microsoft 官方明确把 linear regression、概率统计、hypothesis testing 列为数据科学面试范围Apple 候选人报告直接出现过 linear regression、polynomial regression、covariance 与 NumPy 拟合Uber 公开面经中出现过 “what is logistic regression” 和从零实现 gradient descent 拟合直线Google / Meta 的 MLE 面试则通常把这类基础知识放进 ML fundamentals / ML design / coding / modeling 环节。Linear Regression 不能只准备“公式是什么”而要覆盖数学推导、假设、统计解释、feature interaction、regularization、coding、numerical stability、debugging、scaling、production。Q1. What is Linear Regression?Interview AnswerLinear regression models theconditional mean of a targetas alinear function of the input features. In ordinary least squares, weestimate the coefficientsby minimizing the sum of squared residuals between the observed targets and model predictions.The key advantages areinterpretability,computational efficiency,and strong statistical theory.Its main limitations are thelinearity assumption, sensitivity to outliers, and instability under strong multicollinearity.Linear regression estimates the conditional expectation of the target given the input, assuming the conditional mean is correctly specified. Its parameters can be estimated either through a closed-form least-squares solution or iterative optimization such as gradient descent. Importantly, linearity refers to the model being linear in its parameters, not necessarily in the original input features. By applying nonlinear feature transformations, such as polynomial or sinusoidal functions, linear regression can represent nonlinear relationships while remaining linear in its coefficients.“Linear regression offers three major advantages: interpretability, computational efficiency, and a strong statistical foundation. Its coefficients provide a clear description of conditional linear relationships, and its estimation properties are well understood under classical assumptions.Linear Regression 中的 Linear严格来说是指模型对未知参数coefficients / parameters是线性的而不要求模型对原始输入变量 \(x\) 是线性的。是在模型条件假设正确的条件下预测E[Y/X],Linear Regression 的目标是估计给定 \(Xx\) 时 \(Y\) 的条件期望Conditional Expectation模型的条件均值假设正确Correct Conditional Mean Specification是指模型假设的 \(E[Y\mid Xx]\) 的函数形式与真实数据生成过程中的条件期望一致。Linear in parameters线性回归的核心结构。Nonlinear in input features完全允许。Nonlinear in parameters一般属于非线性回归。Linear Regression 为什么既有解析解又可以使用 Gradient Descent两者只是优化同一个目标函数的不同方法。However, linear regression can underfit nonlinear relationships, ordinary least squares is sensitive to influential observations, and strong multicollinearity can make coefficient estimates unstable. In practice, I would use residual diagnostics, cross-validation, conditioning analysis, and regularization to evaluate whether linear regression is appropriate for the problem.”真实业务场景说明 Linear Regression 的三个核心 limitationsLinearity Assumption线性假设、Sensitivity to Outliers对异常值敏感、Multicollinearity多重共线性Feature Engineering为什么手工指定非线性特征不够加性模型意味着每个特征对预测结果的贡献可以单独计算。但是十几种存在 Feature Interaction特征交互Linearity主要关注 model bias / misspecification。Outliers主要关注 robustness / influential observations。Multicollinearity主要关注 estimation variance / identifiability。还需要记住Multicollinearity 不一定降低预测准确率Outlier 不一定是错误数据非线性关系也不一定要求放弃 Linear Regression。Linear Regression 中的 Linear严格来说是指模型对未知参数coefficients / parameters是线性的而不要求模型对原始输入变量 \(x\) 是线性的。Q2. If linear regression can model nonlinear relationships using polynomial features, why do we still need nonlinear models?这道题表面上是在考 Linear Regression实际上涉及六个核心知识点Feature Engineering为什么手工构造特征有局限问题并不是线性回归能不能拟合曲线而是当真实关系很复杂、输入维度很高或者不同特征存在大量非线性交互时手工构建足够好的特征表示是否仍然高效、稳定、可泛化而神经网络等模型可以通过训练改变内部特征表示使可学习的函数族更加灵活。Feature Engineering为什么手工指定非线性特征不够特征可以手工设计但随着业务复杂度增加组合数量迅速增长。Dimensionality为什么多项式特征会导致维度爆炸非线性模型并非一定能够避免所有高维计算问题。例如某些深层神经网络也可能需要大量参数和训练数据。它们的优势在于可以通过参数共享、层次化表示和架构设计在不显式枚举所有多项式特征的情况下表示复杂函数。Overfitting为什么高阶多项式容易过拟合这里最重要的细节是高阶多项式并不必然过拟合。 是否过拟合还取决于样本量、训练数据分布、正则化强度、特征缩放和具体拟合算法。AWS 官方 ML 文档也明确区分了训练与评估数据上的 underfitting 和 overfitting并把调整模型灵活性、特征组合及正则化列为处理方法Inductive Bias不同模型如何假设和学习数据结构不完全是。Ill-conditioning病态矩阵是条件数很大很大很大不等于 Rank Deficiency秩亏是条件数是无穷大了但当特征高度相关时矩阵可能从 ill-conditioned 发展为 rank-deficient。Model Capacity vs. Generalization表达能力强是否意味着预测能力强Production Trade-offs如何在准确率、可解释性、延迟和资源成本之间选型高阶多项式特征可能造成设计矩阵列之间近似线性相关使最小奇异值很小、条件数很大只有当列之间存在精确线性依赖时才是真正的 rank deficiency。Q3. Ridge Regression 为什么能降低 Variance、提高参数稳定性通过 L2 Regularization 惩罚过大的参数使模型不再过度依赖数据中难以确定的参数方向从而牺牲一部分无偏性Unbiasedness换取更低的估计方差Variance。这与前面讨论的 Multicollinearity、Ill-conditioning、Rank Deficiency 直接相关。我们从三个层面理解优化角度 为什么加入 \(\lambda\|\beta\|_2^2\) 后模型更稳定线性代数角度 为什么 Ridge 能解决 \(X^TX\) 奇异或者接近奇异的问题统计角度 为什么 Ridge 引入 Bias却可能降低最终预测 MSERidge 会倾向于选择参数较小的 C而不是允许参数出现大幅度的正负抵消。这是 Ridge 改善参数稳定性的第一个直觉在预测效果相近的参数组合中倾向于选择系数较小的解。但是当特征高度相关时\(X^TX\) 可能接近奇异当特征完全线性相关时它会奇异。Ridge 通过提高最小特征值显著改善了正则化矩阵的条件数。但注意这不意味着原始设计矩阵 \(X\) 本身变得 Full Rank。Ridge 改变的是要求解的优化问题不是修复原始数据的秩。Ridge Regression 相当于使用零均值各向同性 Gaussian Prior 得到 MAP Estimate最大后验估计。encourage average 可以更精确地表达为Ridge regularization encourages coefficients to shrink toward zero by imposing a zero-mean Gaussian prior.Ridge 使最小 Eigenvalue 远离 0从而缩小最大与最小 Eigenvalue 的相对比例改善条件数。Eigenvalue 越小该方向上的 OLS 参数估计 Variance 越大。Ridge regression introduces a zero-centered Gaussian prior that shrinks model coefficients. Algebraically, the L2 penalty shifts the eigenvalues of \(X^TX\) upward by \(\lambda\), improving the conditioning of the regularized system. Statistically, this shrinkage reduces the variance of coefficient estimates, especially in poorly identified directions, at the cost of introducing bias.增加正则化矩阵的最小 Eigenvalue从而降低参数估计的 Variance”Q4. Multicollinearity增加或删除 Feature 为什么会改变 Linear Regression 的 CoefficientsA teammate accidentally removed or modified one feature in a linear regression model. What happens to the coefficients and evaluation metrics?Before analyzing the impact, Id like to clarify whether the feature was removed or modifiedduring training or only during inference,and whether the model wasretrained afterward.The effects on coefficients and evaluation metrics are different in these cases.-------------------------------------------------------------------------------------------------------------------------------首先区分两个概念区分 Model Misspecification模型设定错误 和 Coefficient Instability系数不稳定。1. Nonlinearity非线性问题定义 真实的条件期望函数 \(E[Y\mid X]\) 无法被模型所选择的特征的线性组合充分表达导致模型设定错误Model Misspecification可能产生系统性预测偏差。2. Multicollinearity多重共线性定义 多个输入特征之间存在精确或近似的线性依赖使部分参数的独立贡献难以识别并可能导致系数估计方差增大、数值条件变差或参数不唯一。Nonlinearity:The selected feature representation is insufficient to capture the true conditional mean relationship between the inputs and the target, potentially leading to model misspecification and systematic prediction bias.Multicollinearity:Two or more predictors exhibit strong linear dependencies, making their individual coefficients difficult to estimate reliably. This can increase estimation variance and numerical instability without necessarily degrading predictive accuracy.最关键的一句话Nonlinearity 是 representation problem表示问题Multicollinearity 是 identifiability and estimation stability problem参数可识别性与估计稳定性问题。不同概念对结果的影响多元线性回归中的单个系数本质上利用的是该 Feature 无法被其他 Feature 解释的剩余变化Question 1If I add one feature to linear regression, will the coefficients of existing features change?They can change if the new feature is correlated with existing features. OLS estimates conditional linear relationships, so the conditioning set matters. But adding an orthogonal feature leaves existing OLS coefficients unchanged when the same dataset is used.Question 2If I remove a feature, will the remaining coefficients become biased?Not necessarily. Omitted-variable bias arises under the relevant model assumptions when the removed feature contributes to the outcome and is correlated with remaining predictors. For predictive modeling, bias must also be distinguished from the best linear projection coefficients.Question 3Can multicollinearity produce unstable coefficients but accurate predictions?Yes. Strongly correlated predictors can make individual effects difficult to identify even when their combined contribution is estimated reliably. The stability of predictions also depends on the test distribution.Question 4Why cant we simply remove the feature with the highest VIF?VIF measures coefficient variance inflation due to linear dependence, not predictive usefulness. Removing an important feature can worsen model specification, prediction, or confounding control.Question 5How would you investigate a coefficient that changes sign after a new feature is introduced?I would check feature definitions, data leakage, sample consistency, correlations, VIF, coefficient confidence intervals, regularization sensitivity, and out-of-sample prediction changes before drawing conclusions.In linear regression, adding or removing a feature can change the coefficients of existing features because regression coefficients represent conditional linear relationships.When a newly added feature is correlated with existing predictors and provides additional information about the target, it can substantially change their estimated coefficients. This may reflect the correction of omitted-variable bias rather than a problem with the model.However, if predictors are highly collinear, their individual contributions become difficult to identify. This can increase coefficient variance and make estimates unstable, even when predictive performance remains relatively stable.In production, I would evaluate feature correlations, VIF, matrix conditioning, coefficient uncertainty, and stability across data samples. I would then perform controlled feature ablation and compare out-of-sample performance, interpretability, and operational costs.Depending on the objective, I might use Ridge regularization, remove redundant features, or redesign the feature representation. If the goal is causal inference rather than prediction, I would also examine confounding and identification assumptions.最后把三个容易混淆的概念对应起来概念核心问题典型数学依据Omitted Variable Bias删掉的 Feature 是否使剩余系数产生系统性偏差\(\beta_2\operatorname{Cov}(X_1,X_2)/\operatorname{Var}(X_1)\)Multicollinearity能否稳定地区分相关 Feature 的独立贡献\(VIF_j1/(1-R_j^2)\)Ridge Regularization能否通过适当偏差降低参数估计方差\((X^TX\lambda I)^{-1}X^Ty\)对于 Senior / Staff ML Engineer最关键的判断是Coefficient Change 不等于 Model DegradationCoefficient Stability 也不等于 Model Correctness。生产环境中是否保留特征应由预测目标、泛化能力、估计稳定性、解释要求和实际业务成本共同决定。增加或删除 Feature 后确实可能改变 Coefficients、预测结果以及模型性能。但 Coefficients 的变化方向并不能直接决定 Accuracy、Precision、Recall 或 ROC-AUC 的变化方向。首先区分两类任务Regression 预测连续值例如 Sales、Revenue、House Price。主要使用 MAE、MSE、RMSE、\(R^2\)。Classification 预测类别或类别概率例如用户是否点击广告、是否流失。使用 Precision、Recall、F1、Accuracy、ROC-AUC、PR-AUC。如果任务是二分类通常使用 Logistic Regression其决策分数仍然是特征的线性组合因此可以沿用许多有关特征相关性和系数稳定性的分析。对于 Logistic Regression也可能出现类似的系数变化但 Logistic 系数是 conditional log-odds effects且具有 non-collapsibility 性质。因此不能将上面的 OLS 遗漏变量偏差公式原样套用到 Logistic Regression。-----------------------------------------------------------------------------------------------------------------------------首先Linear Regression 通常不直接使用 ROC这里要明确一下什么时候用roc对普通连续值回归问题ROC-AUC 不是适用指标。Scikit-learn 文档也将 LinearRegression 的 \(R^2\) 与分类问题的 ROC-AUC 区分开来。scikit-learn 1.9.1 documentation1因此面试时你应该问Is the model predicting a continuous target, or are we using its output as a score for a binary classification task?如果面试官说用于二分类评分那么可以讨论 ROC。Q5. 如何证明模型能力是否发生变化Step 2建立正确的对照实验至少比较以下模型或输入配置Experiment目的Original Model Correct Features正常基准Original Model Corrupted Features衡量推理输入损坏影响Retrained Model without Feature衡量删除特征后的最佳适应能力Retrained Model with Fixed Feature验证数据修复效果Regularized / Alternative Model验证更稳健的模型方案注意第二组不能重新训练因为我们要单独测量 Serving Data Corruption 对现有模型的影响。第三组则必须重新训练因为这是为了回答如果这个 Feature 以后真的不可用剩余特征能否支持一个可用模型这两个问题不同不能混在一起比较。Step 3检查三类模型能力A. Prediction Quality对于 ETAMAE平均预测偏差的绝对大小。RMSE对较大 ETA 错误更敏感。P90 / P95 Absolute Error检查尾部预测错误。不同 ETA 区间的误差短途与长途是否受到不同影响。Prediction Bias是否系统性高估或低估 ETA。例如MetricBeforeAfter Feature CorruptionMAE2.1 min3.8 minRMSE3.2 min5.6 minMean Error0.2 min-1.7 minP95 Absolute Error7.0 min12.5 min这些数字是示例不是 Lyft 实际系统数据。B. Statistical Stability如果同事修改了训练 Feature 并重新训练还应该检查Coefficient Delta\(\hat\beta_{\text{new}}-\hat\beta_{\text{old}}\)Coefficient Standard ErrorConfidence IntervalsVIFSingular Values / Condition NumberBootstrap Coefficient Stability系数是否显著改变不应仅通过数值大小判断。例如单位变化可以使系数扩大或缩小 1000 倍而预测完全不变。C. Production Reliability即使模型在离线测试集上表现正常仍要检查Missing Feature RateFeature Distribution DriftTraining–Serving SkewP99 Inference LatencyFeature Freshness线上预测异常率特征缺失时的 Fallback BehaviorStep 4做数据切片分析Lyft 这种场景还应该按业务条件分别评估。例如交通特征被损坏后不同时段的影响可能非常不均匀。Slice可能的影响Rush hour交通特征影响可能更明显Off-peak影响可能较小Downtown拥堵变化复杂Highway trips速度估计误差可能积累Short trips固定时间开销占比较大Long trips距离、路线与交通的影响更复杂这些是需要实验验证的假设而不是可以直接断言的结果。如果整体 MAE 只上升 2%但高峰期 P95 Error 上升 30%仍然可能是严重的生产问题。Step 5如何决定是否恢复或移除 Feature我的标准是如果是意外损坏首先优先恢复正确 Feature Pipeline而不是立即重新设计模型。如果这个 Feature 长期不可用再考虑使用可靠的替代特征。训练不依赖该特征的 Fallback Model。使用合理的缺失值处理策略。使用 Ridge 降低共线性相关的系数不稳定性。通过 Feature Ablation 验证该特征是否真的具有增量价值。如果错误已经影响线上服务应该根据实际风险优先执行回滚、流量隔离或降级方案。修复后在 Shadow Evaluation 或受控流量中验证再逐步恢复。Q6. linear regression 解释Why does linear regression minimize squared error?标准回答Ordinary least squares defines the estimator by minimizing the squared residuals. This can be understood geometrically as projecting the target vector onto the column space of the design matrix. Independently, under an i.i.d. Gaussian noise assumption, the same objective also arises from maximum likelihood estimation MLE.这句话非常重要OLS can be motivated geometrically without a Gaussian assumption; Gaussian noise provides a probabilistic interpretation.也就是说 residual\[ ry-\hat y \]与 \(X\) 的所有 columns 正交。所以不要回答Linear regression assumes errors must be Gaussian.更准确Normality of errors is not required for fitting OLS or for unbiasedness under the standard exogeneity conditions. It is mainly useful for exact finite-sample inference and for interpreting OLS as maximum likelihood under Gaussian noise.Why does Linear Regression estimate the conditional mean?Q7. Linear regression KernelKernel的实质Q8. Linear Regression 本身有哪些缺点怎么优化面试官最后这部分通常是在检查你能否从模型限制转向具体解决方案。Limitation原因优化方案Limited nonlinear representation固定特征空间只能表达其线性组合Polynomial Features、Splines、GAM、GBDTMulticollinearity特征之间存在强线性依赖Ridge、特征重设计、相关性分析Sensitivity to outliersOLS 使用 Squared ErrorHuber Regression、Robust Regression、数据质量检查Overfitting with many features高维与有限样本增加估计不确定性Ridge / Lasso、特征选择、交叉验证Model misspecification遗漏重要变量或关系形式错误残差分析、领域特征、交互项Heteroscedasticity误差方差随输入变化Robust Standard Errors、Weighted Least SquaresSensitivity to distribution shift训练和推理数据分布可能不同Drift Monitoring、适时重新训练、稳健验证这里还应强调两点第一误差不是正态分布并不自动导致 OLS 系数有偏。正态性主要影响某些有限样本统计推断。第二Ridge 主要缓解参数不稳定和过拟合不会自动修复 Feature Corruption、数据泄漏或错误的条件均值函数。First, I would clarify whether the feature was removed or modified during training or inference, and whether the model was retrained.If a feature is removed and the linear regression model is retrained, the coefficients of the remaining features may increase, decrease, change sign, or remain unchanged. The effect depends on the correlation between the removed feature and the remaining predictors, as well as its relationship with the target. If the feature carries important information, removing it may introduce omitted-variable bias. If it is highly redundant, prediction performance might remain nearly unchanged.If the feature is modified, I would investigate the type of modification. For example, rescaling a feature consistently may change its coefficient without changing predictions. However, setting a feature to zero or introducing corrupted values may significantly affect predictions. If the modification occurs only during inference, the trained coefficients remain unchanged.For evaluation, since linear regression predicts a continuous target, I would primarily compare MAE, RMSE, R-squared, prediction bias, and errors across important data segments. ROC-AUC would only be appropriate if the model output were used as a binary classification score. AUC may remain unchanged if the modification preserves the ranking between positive and negative examples.In production, I would reproduce the issue using the same evaluation dataset, compare the original and corrupted feature pipelines, and run controlled feature ablation experiments. I would also investigate coefficient stability, feature correlations, missing values, distribution shifts, and training-serving skew.Finally, I would restore the correct feature pipeline if this was an accidental change. If the feature is no longer available or proves redundant, I would consider retraining the model without it, introducing fallback features, or applying regularization. I would make the final decision based on out-of-sample performance, statistical reliability, and operational impact.混淆metrics 使用场景linear regression 是Q9. 在 Binary Classification二分类 场景下feature 变化对各个指标的影响。最重要的是先区分Feature Deletion Retraining 删除特征后重新训练模型。Feature Zeroing Retraining 将特征所有训练值置为 0再重新训练。Feature Zeroing at Inference 模型保持不变只在推理阶段将特征置为 0。这三种情形对 Coefficients、Predicted Probability、Accuracy、Precision、Recall、F1、ROC-AUC 和 Log Loss 的影响不同。Can two logistic regression models have the same ROC-AUC but different accuracy, recall, or log loss?Yes. ROC-AUC measures ranking performance, while threshold-based classification metrics and probabilistic metrics measure different aspects of model quality.Q10. ROC-AUC vs. PR Curve如何评估模型什么情况下指标会改变ROC-AUC 衡量模型区分正负样本的整体排序能力PR Curve 更直接反映模型找出正样本时Precision 与 Recall 之间的权衡。ROC 和 PR 的变化取决于所有样本之间的排序而不是单个系数或概率的变化幅度ROC-AUC and Precision-Recall curves evaluate different aspects of binary classification.ROC-AUC measures the models global ranking ability by comparing true positive rate against false positive rate across thresholds. The PR curve measures precision against recall, which is particularly useful when the positive class is rare.If a feature is removed or zeroed, both curves may change because the model may assign different relative rankings to positive and negative examples. However, if the modification only changes score magnitudes while preserving the full ranking, both ROC-AUC and the PR curve remain unchanged.Another important difference is class prevalence. Under pure prior probability shift, ROC-AUC can remain unchanged while precision and the PR curve change substantially.In production, I would compare ROC-AUC, average precision, calibration, and performance at the actual operating threshold or business constraint. I would also evaluate performance across important segments and use paired statistical comparisons to determine whether any improvement is reliable.最终记住三个判断Q11.Generalization and regularization把bias–variance、overfitting、model capacity、L1/L2、dropout、early stopping、data augmentation、cross-validation、weight decay、double descent串在一起Generalization is the ability of a model to perform well on unseen samples drawn from the target distribution. Training minimizes empirical risk, but what we ultimately care about is expected population risk. The gap between training and test performance reflects generalization error.Q:How do you identify overfitting and underfitting?Bias–Variance Trade-offWhere: -Bias²: Error from wrong assumptions (e.g., assuming linear when data is nonlinear) -Variance: Error from sensitivity to training data fluctuations -Irreducible Error: Noise in data that no model can eliminateRegularization introduces an inductive bias that restricts the effective hypothesis space or discourages overly complex solutions.通过降低模型复杂 来降低variance\[ \text{reduce effective model complexity} \]data leakageData leakage occurs when information that would not be available at prediction time influences model training, feature construction, model selection, or evaluation.Data leakage 的核心是任何在真实预测时本不应该获得的信息如果提前进入了训练、预处理、特征工程、模型选择或评估流程就属于泄露。因此仅仅“模型没有直接看到 test set”还不够StandardScaler、PCA、imputation、feature selection、target encoding 等所有会从数据中学习参数的步骤都必须只在当前 training partition 上 fit再去 transform validation/test交叉验证时这些 preprocessing 也必须在每个 fold 内重新 fit。对于时间序列要避免使用未来信息预测过去对于用户、患者、设备等重复实体要根据任务决定是否按 group split防止同一实体同时出现在 train/test重复查看 test score 并据此调模型也等于把 test 当成 validation造成 test-set contamination。最重要的两条原则是Frequentist 通常把参数 \(\theta\) 看成固定未知值Bayesian 把参数的不确定性也建模为 probability distribution。MLE 面试题课件