import%20marimo%0A%0A__generated_with%20%3D%20%220.25.0%22%0Aapp%20%3D%20marimo.App()%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20marimo%20as%20mo%0A%0A%20%20%20%20return%20(mo%2C)%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20METADATA%20%3D%20%7B%0A%20%20%20%20%20%20%20%20%22id%22%3A%20%22loss_functions%22%2C%0A%20%20%20%20%20%20%20%20%22name%22%3A%20%22Loss%20Functions%22%2C%0A%20%20%20%20%20%20%20%20%22summary%22%3A%20%22Reference%20sheet%20for%20regression%2C%20classification%2C%20segmentation%2C%20and%20metric-learning%20losses%2C%20with%20gradients%20and%20short%20examples.%22%2C%0A%20%20%20%20%20%20%20%20%22topics%22%3A%20%5B%22loss_functions%22%2C%20%22optimization%22%2C%20%22regression%22%2C%20%22classification%22%5D%2C%0A%20%20%20%20%7D%0A%20%20%20%20mo.md(f%22%23%20Cheat%20Sheet%20%E2%80%94%20%7BMETADATA%5B'name'%5D%7D%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20How%20to%20read%20this%20sheet%0A%0A%20%20%20%20A%20loss%20gives%20one%20number%20that%20measures%20prediction%20error.%20Unless%20a%20section%20says%0A%20%20%20%20otherwise%2C%20%24L%24%20is%20the%20loss%20for%20one%20example%20and%20a%20dataset%20objective%20is%20the%20mean%20of%0A%20%20%20%20those%20losses.%20A%20derivative%20shows%20the%20direction%20in%20which%20the%20prediction%20or%20logit%0A%20%20%20%20must%20move%20to%20reduce%20the%20loss.%0A%0A%20%20%20%20Choose%20a%20loss%20from%20the%20type%20of%20target%20first%3A%20a%20real%20value%2C%20a%20binary%20label%2C%20a%20class%2C%0A%20%20%20%20a%20mask%2C%20or%20an%20embedding%20relation.%20Then%20check%20its%20assumptions%20and%20sensitivity%20to%0A%20%20%20%20large%20errors.%20For%20binary%20classification%2C%20the%20full%20probability%20derivation%20lives%20in%0A%20%20%20%20%5BFrom%20Bernoulli%20to%20Log%20Loss%5D(%2Fconcepts%2Flog_loss).%0A%0A%20%20%20%20%23%23%20Notations%0A%0A%20%20%20%20%7C%20Symbol%20%7C%20Meaning%20%7C%0A%20%20%20%20%7C---%7C---%7C%0A%20%20%20%20%7C%20%24y%24%20%7C%20target%20%7C%0A%20%20%20%20%7C%20%24%5Chat%20y%24%20%7C%20prediction%20%7C%0A%20%20%20%20%7C%20%24r%3D%5Chat%20y-y%24%20%7C%20residual%20%7C%0A%20%20%20%20%7C%20%24z%24%20%7C%20logit%20%7C%0A%20%20%20%20%7C%20%24p%24%20%7C%20predicted%20probability%20%7C%0A%20%20%20%20%7C%20%24N%24%20%7C%20number%20of%20examples%20%7C%0A%20%20%20%20%7C%20%24K%24%20%7C%20number%20of%20classes%20%7C%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Regression%0A%0A%20%20%20%20%7C%20Loss%20%7C%20Per-example%20formula%20%7C%20Derivative%20with%20respect%20to%20%24%5Chat%20y%24%20%7C%20Example%3A%20%24y%3D2%24%2C%20%24%5Chat%20y%3D4%24%20%7C%0A%20%20%20%20%7C---%7C---%7C---%7C---%7C%0A%20%20%20%20%7C%20squared%20error%20%7C%20%24r%5E2%24%20%7C%20%242r%24%20%7C%20%24L%3D4%24%2C%20%24L'%3D4%24%20%7C%0A%20%20%20%20%7C%20MSE%20%7C%20%24%5Cdfrac1N%5Csum_i%20r_i%5E2%24%20%7C%20%24%5Cdfrac%7B2r_i%7D%7BN%7D%24%20%7C%20%24r%3D(1%2C-1)%5CRightarrow%20L%3D1%24%20%7C%0A%20%20%20%20%7C%20RMSE%20%7C%20%24%5Csqrt%7B%5Cdfrac1N%5Csum_i%20r_i%5E2%7D%24%20%7C%20%24%5Cdfrac%7Br_i%7D%7BN%5C%2C%5Cmathrm%7BRMSE%7D%7D%24%20%7C%20%24r%3D(1%2C-1)%5CRightarrow%20L%3D1%24%20%7C%0A%20%20%20%20%7C%20absolute%20error%20%7C%20%24%7Cr%7C%24%20%7C%20%24%5Coperatorname%7Bsgn%7D(r)%24%20if%20%24r%5Cne0%24%20%7C%20%24L%3D2%24%2C%20%24L'%3D1%24%20%7C%0A%20%20%20%20%7C%20MAE%20%7C%20%24%5Cdfrac1N%5Csum_i%7Cr_i%7C%24%20%7C%20%24%5Cdfrac%7B%5Coperatorname%7Bsgn%7D(r_i)%7DN%24%20%7C%20%24r%3D(1%2C-3)%5CRightarrow%20L%3D2%24%20%7C%0A%20%20%20%20%7C%20Log-cosh%20%7C%20%24%5Clog(%5Ccosh%20r)%24%20%7C%20%24%5Ctanh%20r%24%20%7C%20%24r%3D0%5CRightarrow%20L%3D0%24%20%7C%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Huber%20loss%0A%0A%20%20%20%20%24%24L_%5Cdelta(r)%3D%0A%20%20%20%20%5Cbegin%7Bcases%7D%0A%20%20%20%20%5Cdfrac12r%5E2%2C%26%7Cr%7C%5Cle%5Cdelta%5C%5C%0A%20%20%20%20%5Cdelta%5Cleft(%7Cr%7C-%5Cdfrac12%5Cdelta%5Cright)%2C%26%7Cr%7C%3E%5Cdelta%0A%20%20%20%20%5Cend%7Bcases%7D%24%24%0A%0A%20%20%20%20%24%24%5Cfrac%7B%5Cpartial%20L_%5Cdelta%7D%7B%5Cpartial%5Chat%20y%7D%3D%0A%20%20%20%20%5Cbegin%7Bcases%7D%0A%20%20%20%20r%2C%26%7Cr%7C%5Cle%5Cdelta%5C%5C%0A%20%20%20%20%5Cdelta%5Coperatorname%7Bsgn%7D(r)%2C%26%7Cr%7C%3E%5Cdelta%0A%20%20%20%20%5Cend%7Bcases%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24%5Cdelta%3D1%24%0A%0A%20%20%20%20%24%24r%3D0.5%5CRightarrow%20L%3D0.125%2C%5Cquad%20L'%3D0.5%24%24%0A%0A%20%20%20%20%24%24r%3D3%5CRightarrow%20L%3D2.5%2C%5Cquad%20L'%3D1%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Quantile%20loss%0A%0A%20%20%20%20For%20quantile%20%24%5Ctau%5Cin(0%2C1)%24%20and%20%24r%3Dy-%5Chat%20y%24%3A%0A%0A%20%20%20%20%24%24L_%5Ctau(r)%3D%0A%20%20%20%20%5Cbegin%7Bcases%7D%0A%20%20%20%20%5Ctau%20r%2C%26r%5Cge0%5C%5C%0A%20%20%20%20(%5Ctau-1)r%2C%26r%3C0%0A%20%20%20%20%5Cend%7Bcases%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24%5Ctau%3D0.9%24%0A%0A%20%20%20%20%24%24r%3D2%5CRightarrow%20L%3D1.8%24%24%0A%0A%20%20%20%20%24%24r%3D-2%5CRightarrow%20L%3D0.2%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Binary%20cross-entropy%0A%0A%20%20%20%20With%20%24y%5Cin%5C%7B0%2C1%5C%7D%24%20and%20%24p%5Cin(0%2C1)%24%3A%0A%0A%20%20%20%20%24%24%5Cboxed%7BL%3D-%5Cleft%5By%5Clog%20p%2B(1-y)%5Clog(1-p)%5Cright%5D%7D%24%24%0A%0A%20%20%20%20%24%24%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20p%7D%3D-%5Cfrac%20yp%2B%5Cfrac%7B1-y%7D%7B1-p%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24y%3D1%24%2C%20%24p%3D0.8%24%0A%0A%20%20%20%20%24%24L%3D-%5Clog(0.8)%5Capprox0.223%24%24%0A%0A%20%20%20%20%23%23%23%20BCE%20with%20logits%0A%0A%20%20%20%20With%20%24p%3D%5Csigma(z)%24%3A%0A%0A%20%20%20%20%24%24%5Cboxed%7B%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20z%7D%3Dp-y%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24y%3D1%24%2C%20%24z%3D0%24%2C%20so%20%24p%3D0.5%24%0A%0A%20%20%20%20%24%24%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20z%7D%3D0.5-1%3D-0.5%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Multiclass%20cross-entropy%0A%0A%20%20%20%20With%20a%20one-hot%20target%20%24y%24%20and%20%24p%3D%5Coperatorname%7Bsoftmax%7D(z)%24%3A%0A%0A%20%20%20%20%24%24%5Cboxed%7BL%3D-%5Csum_%7Bk%3D1%7D%5E%7BK%7Dy_k%5Clog%20p_k%7D%24%24%0A%0A%20%20%20%20If%20the%20correct%20class%20is%20%24c%24%3A%0A%0A%20%20%20%20%24%24L%3D-%5Clog%20p_c%24%24%0A%0A%20%20%20%20Gradient%20with%20respect%20to%20the%20logits%3A%0A%0A%20%20%20%20%24%24%5Cboxed%7B%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20z_k%7D%3Dp_k-y_k%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24y%3D(1%2C0%2C0)%24%2C%20%24p%3D(0.7%2C0.2%2C0.1)%24%0A%0A%20%20%20%20%24%24L%3D-%5Clog(0.7)%5Capprox0.357%24%24%0A%0A%20%20%20%20%24%24%5Cnabla_zL%3Dp-y%3D(-0.3%2C0.2%2C0.1)%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Negative%20log-likelihood%0A%0A%20%20%20%20%24%24%5Cboxed%7BL%3D-%5Clog%20p(y%5Cmid%20x)%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20the%20model%20assigns%20%24p(y%5Cmid%20x)%3D0.9%24%20to%20the%20correct%20class%3A%0A%0A%20%20%20%20%24%24L%3D-%5Clog(0.9)%5Capprox0.105%24%24%0A%0A%20%20%20%20For%20%60log_softmax%60%20followed%20by%20NLL%3A%0A%0A%20%20%20%20%24%24L%3D-%5Coperatorname%7BLogSoftmax%7D(z)_c%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Focal%20loss%0A%0A%20%20%20%20For%20%24p_t%24%2C%20the%20probability%20of%20the%20correct%20class%3A%0A%0A%20%20%20%20%24%24%5Cboxed%7BL%3D-%5Calpha_t(1-p_t)%5E%5Cgamma%5Clog%20p_t%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24%5Calpha_t%3D1%24%2C%20%24%5Cgamma%3D2%24%0A%0A%20%20%20%20%24%24p_t%3D0.9%5CRightarrow%20L%3D-(0.1)%5E2%5Clog(0.9)%5Capprox0.0011%24%24%0A%0A%20%20%20%20%24%24p_t%3D0.2%5CRightarrow%20L%3D-(0.8)%5E2%5Clog(0.2)%5Capprox1.030%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Hinge%20loss%0A%0A%20%20%20%20With%20%24y%5Cin%5C%7B-1%2C%2B1%5C%7D%24%20and%20score%20%24s%24%3A%0A%0A%20%20%20%20%24%24%5Cboxed%7BL%3D%5Cmax(0%2C1-ys)%7D%24%24%0A%0A%20%20%20%20%24%24%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20s%7D%3D%0A%20%20%20%20%5Cbegin%7Bcases%7D%0A%20%20%20%20-y%2C%26ys%3C1%5C%5C%0A%20%20%20%200%2C%26ys%3E1%0A%20%20%20%20%5Cend%7Bcases%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24y%3D1%24%2C%20%24s%3D0.4%24%0A%0A%20%20%20%20%24%24L%3D%5Cmax(0%2C1-0.4)%3D0.6%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20KL%20divergence%0A%0A%20%20%20%20For%20two%20discrete%20distributions%20%24P%24%20and%20%24Q%24%3A%0A%0A%20%20%20%20%24%24%5Cboxed%7BD_%7BKL%7D(P%5C%7CQ)%3D%5Csum_iP_i%5Clog%5Cfrac%7BP_i%7D%7BQ_i%7D%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24P%3D(0.5%2C0.5)%24%2C%20%24Q%3D(0.5%2C0.5)%24%0A%0A%20%20%20%20%24%24D_%7BKL%7D(P%5C%7CQ)%3D0%24%24%0A%0A%20%20%20%20Relation%20to%20cross-entropy%3A%0A%0A%20%20%20%20%24%24H(P%2CQ)%3DH(P)%2BD_%7BKL%7D(P%5C%7CQ)%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Segmentation%0A%0A%20%20%20%20Let%20%24p_i%24%20be%20the%20predictions%20and%20%24y_i%5Cin%5C%7B0%2C1%5C%7D%24%20the%20masks.%0A%0A%20%20%20%20%23%23%23%20Dice%20loss%0A%0A%20%20%20%20%24%24%5Coperatorname%7BDice%7D%3D%5Cfrac%7B2%5Csum_ip_iy_i%2B%5Cvarepsilon%7D%7B%5Csum_ip_i%2B%5Csum_iy_i%2B%5Cvarepsilon%7D%24%24%0A%0A%20%20%20%20%24%24%5Cboxed%7BL_%7BDice%7D%3D1-%5Coperatorname%7BDice%7D%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20perfect%20prediction%20%24p%3Dy%24%0A%0A%20%20%20%20%24%24%5Coperatorname%7BDice%7D%3D1%5CRightarrow%20L_%7BDice%7D%3D0%24%24%0A%0A%20%20%20%20%23%23%23%20IoU%20%2F%20Jaccard%20loss%0A%0A%20%20%20%20%24%24%5Coperatorname%7BIoU%7D%3D%5Cfrac%7B%5Csum_ip_iy_i%2B%5Cvarepsilon%7D%0A%20%20%20%20%7B%5Csum_ip_i%2B%5Csum_iy_i-%5Csum_ip_iy_i%2B%5Cvarepsilon%7D%24%24%0A%0A%20%20%20%20%24%24%5Cboxed%7BL_%7BIoU%7D%3D1-%5Coperatorname%7BIoU%7D%7D%24%24%0A%0A%20%20%20%20%23%23%23%20Tversky%20loss%0A%0A%20%20%20%20%24%24T%3D%5Cfrac%7BTP%2B%5Cvarepsilon%7D%7BTP%2B%5Calpha%20FP%2B%5Cbeta%20FN%2B%5Cvarepsilon%7D%24%24%0A%0A%20%20%20%20%24%24%5Cboxed%7BL_%7BTversky%7D%3D1-T%7D%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Metric%20learning%0A%0A%20%20%20%20%23%23%23%20Contrastive%20loss%0A%0A%20%20%20%20With%20%24y%3D1%24%20for%20a%20similar%20pair%2C%20distance%20%24d%24%2C%20and%20margin%20%24m%24%3A%0A%0A%20%20%20%20%24%24%5Cboxed%7BL%3Dy%5C%2Cd%5E2%2B(1-y)%5Cmax(0%2Cm-d)%5E2%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20similar%20pair%2C%20%24y%3D1%24%2C%20%24d%3D0.3%24%0A%0A%20%20%20%20%24%24L%3D0.3%5E2%3D0.09%24%24%0A%0A%20%20%20%20%23%23%23%20Triplet%20loss%0A%0A%20%20%20%20%24%24%5Cboxed%7BL%3D%5Cmax(0%2Cd(a%2Cp)-d(a%2Cn)%2Bm)%7D%24%24%0A%0A%20%20%20%20**Example%3A**%20%24d(a%2Cp)%3D0.4%24%2C%20%24d(a%2Cn)%3D1.0%24%2C%20%24m%3D0.3%24%0A%0A%20%20%20%20%24%24L%3D%5Cmax(0%2C0.4-1%2B0.3)%3D0%24%24%0A%0A%20%20%20%20%23%23%23%20Cosine%20embedding%20loss%0A%0A%20%20%20%20%24%24%5Ccos(x_1%2Cx_2)%3D%5Cfrac%7Bx_1%5ETx_2%7D%7B%5C%7Cx_1%5C%7C%5C%7Cx_2%5C%7C%7D%24%24%0A%0A%20%20%20%20For%20a%20similar%20pair%3A%0A%0A%20%20%20%20%24%24L%3D1-%5Ccos(x_1%2Cx_2)%24%24%0A%0A%20%20%20%20**Example%3A**%20identical%20vectors%20%24%5CRightarrow%20%5Ccos%3D1%20%5CRightarrow%20L%3D0%24.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Regularization%20added%20to%20the%20loss%0A%0A%20%20%20%20%7C%20Regularization%20%7C%20Formula%20%7C%20Gradient%20%7C%0A%20%20%20%20%7C---%7C---%7C---%7C%0A%20%20%20%20%7C%20%24L_1%24%20%7C%20%24%5Clambda%5Csum_j%7Cw_j%7C%24%20%7C%20%24%5Clambda%5Coperatorname%7Bsgn%7D(w_j)%24%20%7C%0A%20%20%20%20%7C%20%24L_2%24%20%7C%20%24%5Clambda%5Csum_jw_j%5E2%24%20%7C%20%242%5Clambda%20w_j%24%20%7C%0A%20%20%20%20%7C%20%24L_2%24%20with%20factor%20%241%2F2%24%20%7C%20%24%5Cdfrac%5Clambda2%5Csum_jw_j%5E2%24%20%7C%20%24%5Clambda%20w_j%24%20%7C%0A%0A%20%20%20%20**Example%3A**%20%24w%3D(1%2C-2)%24%20and%20%24%5Clambda%3D0.1%24%0A%0A%20%20%20%20%24%24L_1%3D0.1(1%2B2)%3D0.3%24%24%0A%0A%20%20%20%20%24%24L_2%3D0.1(1%5E2%2B(-2)%5E2)%3D0.5%24%24%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Quick%20choice%0A%0A%20%20%20%20%7C%20Task%20%7C%20Common%20loss%20%7C%0A%20%20%20%20%7C---%7C---%7C%0A%20%20%20%20%7C%20regression%20%7C%20MSE%20%7C%0A%20%20%20%20%7C%20regression%20with%20outliers%20%7C%20MAE%20or%20Huber%20%7C%0A%20%20%20%20%7C%20quantile%20prediction%20%7C%20quantile%20loss%20%7C%0A%20%20%20%20%7C%20binary%20classification%20%7C%20BCE%20with%20logits%20%7C%0A%20%20%20%20%7C%20multilabel%20classification%20%7C%20BCE%20with%20logits%20per%20class%20%7C%0A%20%20%20%20%7C%20multiclass%20classification%20%7C%20cross-entropy%20with%20logits%20%7C%0A%20%20%20%20%7C%20highly%20imbalanced%20classes%20%7C%20weighted%20cross-entropy%20or%20focal%20loss%20%7C%0A%20%20%20%20%7C%20segmentation%20%7C%20BCE%2FCE%20%2B%20Dice%20%7C%0A%20%20%20%20%7C%20embeddings%20%7C%20contrastive%20or%20triplet%20loss%20%7C%0A%0A%20%20%20%20%23%23%20Related%20resources%0A%0A%20%20%20%20-%20%5BFrom%20Bernoulli%20to%20Log%20Loss%5D(%2Fconcepts%2Flog_loss)%20%E2%80%94%20full%20derivation%20of%20binary%20cross-entropy.%0A%20%20%20%20-%20%5BLog%20Loss%20Formula%20Sheet%5D(%2Fcheatsheets%2Flog_loss)%20%E2%80%94%20compact%20reference%20for%20binary%20classification.%0A%20%20%20%20-%20%5BLoss%20and%20Optimization%5D(%2Fconcepts%2Floss_optimization)%20%E2%80%94%20the%20role%20of%20a%20loss%20during%20training.%0A%20%20%20%20-%20%5BDeep%20Learning%20Essential%20Formulas%5D(%2Fcheatsheets%2Fdeep_learning)%20%E2%80%94%20where%20the%20loss%20fits%20in%20the%20full%20training%20loop.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
2ccf3b7de42ce5eddc6e427ef76f75cd