import%20marimo%0A%0A__generated_with%20%3D%20%220.25.0%22%0Aapp%20%3D%20marimo.App()%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20marimo%20as%20mo%0A%0A%20%20%20%20return%20(mo%2C)%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20METADATA%20%3D%20%7B%0A%20%20%20%20%20%20%20%20%22id%22%3A%20%22bias_variance%22%2C%0A%20%20%20%20%20%20%20%20%22name%22%3A%20%22Bias%20and%20Variance%22%2C%0A%20%20%20%20%20%20%20%20%22kind%22%3A%20%22concept%22%2C%0A%20%20%20%20%20%20%20%20%22difficulty%22%3A%20%22intermediate%22%2C%0A%20%20%20%20%20%20%20%20%22status%22%3A%20%22complete%22%2C%0A%20%20%20%20%20%20%20%20%22topics%22%3A%20%5B%22generalization%22%2C%20%22model_complexity%22%5D%2C%0A%20%20%20%20%20%20%20%20%22related_models%22%3A%20%5B%5D%2C%0A%20%20%20%20%20%20%20%20%22prerequisites%22%3A%20%5B%22generalization%22%5D%2C%0A%20%20%20%20%7D%0A%20%20%20%20mo.md(f%22%23%20%7BMETADATA%5B'name'%5D%7D%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20The%20central%20tension%0A%0A%20%20%20%20Bias%20is%20systematic%20error%20from%20an%20overly%20restrictive%20model%3B%20variance%20is%20instability%20caused%20by%20sensitivity%20to%20the%20particular%20training%20sample.%0A%0A%20%20%20%20%23%23%20Imagine%20many%20possible%20training%20sets%0A%0A%20%20%20%20Refit%20the%20same%20modeling%20procedure%20on%20many%20equally%20plausible%20datasets.%20If%20all%20predictions%20miss%20in%20the%20same%20direction%2C%20suspect%20bias.%20If%20predictions%20move%20dramatically%20between%20samples%2C%20suspect%20variance.%0A%0A%20%20%20%20%23%23%20Why%20one%20fitted%20dataset%20is%20not%20enough%0A%0A%20%20%20%20The%20two%20failure%20patterns%20suggest%20different%20responses%3A%20more%20flexibility%20can%20reduce%20bias%2C%20while%20regularization%2C%20averaging%2C%20or%20more%20representative%20data%20can%20reduce%20variance.%0A%0A%20%20%20%20%23%23%20What%20we%20will%20vary%0A%0A%20%20%20%20**Question%3A**%20how%20do%20polynomial%20models%20behave%20when%20the%20training%20sample%20is%20repeatedly%20redrawn%3F%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20One%20dataset%20is%20only%20one%20possible%20sample%0A%0A%20%20%20%20Imagine%20repeatedly%20collecting%20a%20new%20training%20set%20from%20the%20same%20conditions.%20Apply%20the%20same%20learning%20procedure%20to%20each%20set.%20This%20produces%20fitted%20functions%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Chat%20f%5E%7B(1)%7D(x)%2C%5Chat%20f%5E%7B(2)%7D(x)%2C%5Cldots%0A%20%20%20%20%24%24%0A%0A%20%20%20%20At%20one%20input%20%24x%24%2C%20their%20average%20prediction%20is%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathbb%20E_D%5B%5Chat%20f_D(x)%5D%2C%0A%20%20%20%20%24%24%0A%0A%20%20%20%20where%20%24D%24%20represents%20the%20random%20training%20set.%0A%0A%20%20%20%20%23%23%20Bias%0A%0A%20%20%20%20Bias%20measures%20the%20difference%20between%20the%20average%20fitted%20prediction%20and%20the%20true%20signal%20%24f(x)%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cboxed%7B%0A%20%20%20%20%5Coperatorname%7BBias%7D(x)%0A%20%20%20%20%3D%5Cmathbb%20E_D%5B%5Chat%20f_D(x)%5D-f(x)%0A%20%20%20%20%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20High%20bias%20means%20that%20even%20after%20averaging%20many%20fitted%20versions%2C%20the%20procedure%20misses%20the%20same%20structure%20in%20the%20same%20way.%20A%20straight%20line%20fitted%20to%20a%20strongly%20curved%20signal%20is%20a%20simple%20example.%0A%0A%20%20%20%20%23%23%20Variance%0A%0A%20%20%20%20Variance%20measures%20how%20much%20fitted%20predictions%20move%20when%20the%20training%20sample%20changes%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cboxed%7B%0A%20%20%20%20%5Coperatorname%7BVar%7D(x)%0A%20%20%20%20%3D%5Cmathbb%20E_D%5Cleft%5B%0A%20%20%20%20(%5Chat%20f_D(x)-%5Cmathbb%20E_D%5B%5Chat%20f_D(x)%5D)%5E2%0A%20%20%20%20%5Cright%5D%0A%20%20%20%20%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20High%20variance%20means%20that%20small%20changes%20in%20the%20data%20create%20large%20changes%20in%20the%20fitted%20rule.%0A%0A%20%20%20%20%23%23%20The%20squared-error%20decomposition%0A%0A%20%20%20%20Assume%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20Y%3Df(X)%2B%5Cvarepsilon%2C%0A%20%20%20%20%5Cqquad%0A%20%20%20%20%5Cmathbb%20E%5B%5Cvarepsilon%5D%3D0%2C%0A%20%20%20%20%5Cqquad%0A%20%20%20%20%5Coperatorname%7BVar%7D(%5Cvarepsilon)%3D%5Csigma%5E2.%0A%20%20%20%20%24%24%0A%0A%20%20%20%20At%20a%20fixed%20input%20%24x%24%2C%20expected%20squared%20prediction%20error%20can%20be%20decomposed%20as%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cboxed%7B%0A%20%20%20%20%5Cmathbb%20E%5B(Y-%5Chat%20f_D(x))%5E2%5D%0A%20%20%20%20%3D%5Csigma%5E2%0A%20%20%20%20%2B%5Coperatorname%7BBias%7D(x)%5E2%0A%20%20%20%20%2B%5Coperatorname%7BVar%7D(x)%0A%20%20%20%20%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20The%20three%20terms%20have%20different%20meanings%3A%0A%0A%20%20%20%20%7C%20Term%20%7C%20Source%20%7C%20Can%20fitting%20remove%20it%3F%20%7C%0A%20%20%20%20%7C---%7C---%7C---%7C%0A%20%20%20%20%7C%20%24%5Csigma%5E2%24%20%7C%20noise%20in%20the%20outcome%20%7C%20no%2C%20not%20from%20the%20available%20input%20alone%20%7C%0A%20%20%20%20%7C%20bias%24%5E2%24%20%7C%20systematic%20restriction%20of%20the%20procedure%20%7C%20sometimes%2C%20with%20useful%20flexibility%20or%20features%20%7C%0A%20%20%20%20%7C%20variance%20%7C%20sensitivity%20to%20the%20sampled%20data%20%7C%20sometimes%2C%20with%20more%20data%20or%20stronger%20constraints%20%7C%0A%0A%20%20%20%20%23%23%20A%20short%20derivation%0A%0A%20%20%20%20Write%20the%20error%20around%20the%20average%20prediction%20%24%5Cbar%20f(x)%3D%5Cmathbb%20E_D%5B%5Chat%20f_D(x)%5D%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20Y-%5Chat%20f_D(x)%0A%20%20%20%20%3D%5Cvarepsilon%2B%5Bf(x)-%5Cbar%20f(x)%5D%2B%5B%5Cbar%20f(x)-%5Chat%20f_D(x)%5D.%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Square%20both%20sides%20and%20take%20expectations.%20The%20cross%20terms%20vanish%20because%20the%20noise%20has%20mean%20zero%20and%20the%20fitted%20deviations%20average%20to%20zero.%20The%20remaining%20three%20squared%20terms%20are%20noise%2C%20bias%20squared%2C%20and%20variance.%0A%0A%20%20%20%20%23%23%20Reading%20common%20symptoms%0A%0A%20%20%20%20-%20High%20training%20and%20validation%20error%20can%20suggest%20high%20bias.%0A%20%20%20%20-%20Very%20low%20training%20error%20with%20much%20higher%20validation%20error%20can%20suggest%20high%20variance.%0A%20%20%20%20-%20These%20are%20clues%2C%20not%20proofs.%20Leakage%2C%20distribution%20shift%2C%20poor%20optimization%2C%20and%20wrong%20labels%20can%20create%20similar%20patterns.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20numpy%20as%20np%0A%20%20%20%20import%20matplotlib.pyplot%20as%20plt%0A%20%20%20%20from%20sklearn.linear_model%20import%20LinearRegression%0A%20%20%20%20from%20sklearn.pipeline%20import%20make_pipeline%0A%20%20%20%20from%20sklearn.preprocessing%20import%20PolynomialFeatures%0A%0A%20%20%20%20rng%20%3D%20np.random.default_rng(4)%0A%20%20%20%20grid%20%3D%20np.linspace(-2%2C%202%2C%20160)%0A%20%20%20%20truth%20%3D%20np.sin(2%20*%20grid)%0A%20%20%20%20degrees%20%3D%20%5B1%2C%205%2C%2014%5D%0A%20%20%20%20predictions%20%3D%20%7Bdegree%3A%20%5B%5D%20for%20degree%20in%20degrees%7D%0A%20%20%20%20for%20_%20in%20range(45)%3A%0A%20%20%20%20%20%20%20%20x%20%3D%20np.sort(rng.uniform(-2%2C%202%2C%2024))%0A%20%20%20%20%20%20%20%20y%20%3D%20np.sin(2%20*%20x)%20%2B%20rng.normal(0%2C%200.28%2C%20len(x))%0A%20%20%20%20%20%20%20%20for%20fit_degree%20in%20degrees%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20model%20%3D%20make_pipeline(PolynomialFeatures(fit_degree)%2C%20LinearRegression())%0A%20%20%20%20%20%20%20%20%20%20%20%20model.fit(x%5B%3A%2C%20None%5D%2C%20y)%0A%20%20%20%20%20%20%20%20%20%20%20%20predictions%5Bfit_degree%5D.append(model.predict(grid%5B%3A%2C%20None%5D))%0A%20%20%20%20return%20degrees%2C%20grid%2C%20np%2C%20plt%2C%20predictions%2C%20truth%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20The%20code%20above%20repeatedly%20draws%20a%20new%20training%20sample%20from%20the%20same%20process.%20For%0A%20%20%20%20each%20polynomial%20degree%2C%20we%20now%20plot%20the%20mean%20prediction%20and%20the%20variation%20between%0A%20%20%20%20fits.%20This%20makes%20bias%20and%20variance%20visible%20as%20two%20different%20geometric%20patterns.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(degrees%2C%20grid%2C%20np%2C%20plt%2C%20predictions%2C%20truth)%3A%0A%20%20%20%20fig%2C%20axes%20%3D%20plt.subplots(1%2C%203%2C%20figsize%3D(13%2C%203.8)%2C%20sharey%3DTrue)%0A%20%20%20%20for%20ax%2C%20display_degree%20in%20zip(axes%2C%20degrees)%3A%0A%20%20%20%20%20%20%20%20draws%20%3D%20np.asarray(predictions%5Bdisplay_degree%5D)%0A%20%20%20%20%20%20%20%20mean_prediction%20%3D%20draws.mean(axis%3D0)%0A%20%20%20%20%20%20%20%20spread%20%3D%20draws.std(axis%3D0)%0A%20%20%20%20%20%20%20%20bias_squared%20%3D%20np.mean((mean_prediction%20-%20truth)%20**%202)%0A%20%20%20%20%20%20%20%20variance%20%3D%20np.mean(spread%20**%202)%0A%20%20%20%20%20%20%20%20ax.plot(grid%2C%20truth%2C%20color%3D%22black%22%2C%20linestyle%3D%22--%22%2C%20label%3D%22true%20signal%22)%0A%20%20%20%20%20%20%20%20ax.plot(grid%2C%20mean_prediction%2C%20color%3D%22crimson%22%2C%20label%3D%22mean%20prediction%22)%0A%20%20%20%20%20%20%20%20ax.fill_between(grid%2C%20mean_prediction%20-%20spread%2C%20mean_prediction%20%2B%20spread%2C%20alpha%3D0.25%2C%20label%3D%22sample%20instability%22)%0A%20%20%20%20%20%20%20%20ax.set(title%3Df%22degree%20%7Bdisplay_degree%7D%5Cnbias%C2%B2%20%7Bbias_squared%3A.3f%7D%20%C2%B7%20variance%20%7Bvariance%3A.3f%7D%22%2C%20xlabel%3D%22x%22)%0A%20%20%20%20%20%20%20%20ax.set_ylim(-2%2C%202)%0A%20%20%20%20axes%5B0%5D.set_ylabel(%22prediction%22)%0A%20%20%20%20axes%5B0%5D.legend(fontsize%3D8)%0A%20%20%20%20plt.tight_layout()%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20A%20misleading%20diagnosis%0A%0A%20%20%20%20Diagnosing%20every%20validation%20failure%20as%20variance.%20Leakage%2C%20poor%20features%2C%20distribution%20shift%2C%20optimization%20failure%2C%20and%20label%20noise%20can%20produce%20similar%20symptoms.%0A%0A%20%20%20%20%23%23%20Read%20the%20pattern%20across%20samples%0A%0A%20%20%20%20Compare%20training%20and%20validation%20errors%2C%20then%20refit%20across%20resamples.%20Persistent%20error%20suggests%20bias%3B%20large%20movement%20between%20fits%20suggests%20variance.%0A%0A%20%20%20%20%23%23%20Decision%20rule%0A%0A%20%20%20%20Increase%20flexibility%20only%20for%20demonstrated%20underfitting%3B%20use%20regularization%2C%20averaging%2C%20or%20better%20data%20when%20the%20procedure%20is%20unstable%20across%20samples.%0A%0A%20%20%20%20%23%23%20Previous%20and%20next%0A%0A%20%20%20%20-%20Previous%3A%20%5BBackpropagation%5D(%2Fconcepts%2Fbackpropagation)%20explains%20how%20complex%20fitted%20rules%20receive%20gradients.%0A%20%20%20%20-%20Next%3A%20%5BRegularization%5D(%2Fconcepts%2Fregularization)%20shows%20one%20way%20to%20reduce%20unstable%20fitting.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
1f17bc17dd986b1987be8ee3bc173b80