import%20marimo%0A%0A__generated_with%20%3D%20%220.25.0%22%0Aapp%20%3D%20marimo.App()%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20marimo%20as%20mo%0A%0A%20%20%20%20return%20(mo%2C)%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20METADATA%20%3D%20%7B%0A%20%20%20%20%20%20%20%20%22id%22%3A%20%22tcn%22%2C%0A%20%20%20%20%20%20%20%20%22name%22%3A%20%22Temporal%20Convolutional%20Network%22%2C%0A%20%20%20%20%20%20%20%20%22types%22%3A%20%5B%22architecture%22%5D%2C%0A%20%20%20%20%20%20%20%20%22families%22%3A%20%5B%22neural_networks%22%2C%20%22convolutional%22%2C%20%22sequential%22%5D%2C%0A%20%20%20%20%20%20%20%20%22tasks%22%3A%20%5B%22classification%22%2C%20%22regression%22%2C%20%22forecasting%22%2C%20%22representation%22%5D%2C%0A%20%20%20%20%20%20%20%20%22data%22%3A%20%5B%22time_series%22%2C%20%22sequences%22%2C%20%22audio%22%5D%2C%0A%20%20%20%20%20%20%20%20%22learning%22%3A%20%5B%22supervised%22%2C%20%22self_supervised%22%5D%2C%0A%20%20%20%20%20%20%20%20%22capacity%22%3A%20%22parametric%22%2C%0A%20%20%20%20%20%20%20%20%22mechanisms%22%3A%20%5B%22causal_convolution%22%2C%20%22dilation%22%2C%20%22backpropagation%22%5D%2C%0A%20%20%20%20%20%20%20%20%22properties%22%3A%20%5B%22nonlinear%22%2C%20%22representation_learning%22%5D%2C%0A%20%20%20%20%20%20%20%20%22constraints%22%3A%20%5B%22requires_large_data%22%2C%20%22requires_scaling%22%2C%20%22sensitive_to_tuning%22%5D%2C%0A%20%20%20%20%20%20%20%20%22difficulty%22%3A%20%22intermediate%22%2C%0A%20%20%20%20%20%20%20%20%22status%22%3A%20%22complete%22%2C%0A%20%20%20%20%20%20%20%20%22explainability%22%3A%20%22low%22%2C%0A%20%20%20%20%20%20%20%20%22training_cost%22%3A%20%22high%22%2C%0A%20%20%20%20%20%20%20%20%22inference_cost%22%3A%20%22medium%22%2C%0A%20%20%20%20%20%20%20%20%22data_appetite%22%3A%20%22high%22%2C%0A%20%20%20%20%7D%0A%0A%20%20%20%20mo.md(f%22%23%20%7BMETADATA%5B'name'%5D%7D%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20The%20sequence%20rule%0A%0A%20%20%20%20A%20TCN%20uses%20causal%2C%20usually%20dilated%20one-dimensional%20convolutions%20to%20model%20sequences%20with%20a%20large%20and%20controllable%20history.%0A%0A%20%20%20%20%23%23%20Build%20causal%20memory%20with%20spaced%20connections%0A%0A%20%20%20%20Instead%20of%20reading%20a%20sequence%20one%20step%20at%20a%20time%2C%20a%20TCN%20slides%20pattern%20detectors%20across%20many%20positions%20in%20parallel.%20Deeper%20layers%20combine%20short%20motifs%20into%20longer%20patterns.%0A%0A%20%20%20%20Dilation%20creates%20gaps%20between%20sampled%20positions.%20A%20layer%20with%20dilation%201%20sees%20nearby%20values%3B%20later%20layers%20with%20dilation%202%2C%204%2C%20and%208%20reach%20farther%20back%20without%20needing%20huge%20kernels.%0A%0A%20%20%20%20%23%23%20Why%20it%20belongs%20to%20several%20categories%0A%0A%20%20%20%20TCN%20is%20a%20neural-network%20architecture%20because%20it%20learns%20layered%20representations%3B%20convolutional%20because%20its%20main%20operation%20is%201D%20convolution%3B%20sequential%20because%20it%20models%20ordered%20data%3B%20and%20compatible%20with%20supervised%20or%20self-supervised%20objectives.%0A%0A%20%20%20%20%23%23%20Input%20and%20output%0A%0A%20%20%20%20-%20**Input%3A**%20sequences%20shaped%20as%20time%20steps%20by%20channels%20or%20features.%0A%20%20%20%20-%20**Output%3A**%20one%20prediction%20per%20sequence%2C%20one%20per%20time%20step%2C%20or%20a%20future%20horizon.%0A%20%20%20%20-%20**Learns%3A**%20temporal%20convolution%20filters%20and%20usually%20residual%20transformations.%0A%0A%20%20%20%20%23%23%20Core%20ingredients%0A%0A%20%20%20%20**Causal%20convolution**%20%E2%80%94%20the%20output%20at%20time%20%24t%24%20depends%20only%20on%20time%20%24t%24%20and%20earlier%20positions.%20Padding%20must%20be%20designed%20carefully%3B%20ordinary%20symmetric%20padding%20can%20leak%20future%20information.%0A%0A%20%20%20%20**Dilation**%20%E2%80%94%20dilated%20kernels%20expand%20context%20efficiently.%20For%20kernel%20size%20%24k%24%20and%20dilations%20%24d_l%24%2C%20the%20receptive%20field%20of%20a%20simple%20stack%20is%3A%0A%0A%20%20%20%20%24%24R%20%3D%201%20%2B%20%5Csum_l%20(k-1)d_l%24%24%0A%0A%20%20%20%20**Residual%20blocks**%20%E2%80%94%20skip%20connections%20help%20gradients%20and%20let%20deeper%20stacks%20refine%20rather%20than%20completely%20replace%20representations.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20Step%20by%20step%3A%20one%20causal%20dilated%20convolution%0A%0A%20%20%20%20For%20kernel%20weights%20%24w_0%2C%5Cldots%2Cw_%7Bk-1%7D%24%20and%20dilation%20%24d%24%2C%20a%20causal%20convolution%20at%20time%20%24t%24%20is%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cboxed%7B%0A%20%20%20%20y_t%3D%5Csum_%7Bj%3D0%7D%5E%7Bk-1%7Dw_jx_%7Bt-jd%7D%0A%20%20%20%20%7D%0A%20%20%20%20%24%24%0A%0A%20%20%20%20Only%20indices%20at%20or%20before%20%24t%24%20appear.%20This%20is%20what%20makes%20the%20operation%20causal.%0A%0A%20%20%20%20Let%20%24k%3D3%24%2C%20%24d%3D2%24%2C%20and%20%24t%3D6%24.%20The%20output%20uses%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20x_6%2C%5Cquad%20x_4%2C%5Cquad%20x_2.%0A%20%20%20%20%24%24%0A%0A%20%20%20%20It%20cannot%20use%20%24x_7%24%20or%20any%20later%20value.%20With%20weights%20%24(0.5%2C0.3%2C-0.2)%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20y_6%3D0.5x_6%2B0.3x_4-0.2x_2.%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%23%23%20Build%20a%20receptive%20field%0A%0A%20%20%20%20One%20layer%20with%20kernel%20size%20%24k%24%20and%20dilation%20%24d%24%20reaches%20back%20%24(k-1)d%24%20steps.%20For%20a%20simple%20stack%20with%20dilations%20%24d_1%2C%5Cldots%2Cd_L%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20R%3D1%2B%5Csum_%7Bl%3D1%7D%5E%7BL%7D(k-1)d_l.%0A%20%20%20%20%24%24%0A%0A%20%20%20%20With%20%24k%3D3%24%20and%20dilations%20%24(1%2C2%2C4%2C8)%24%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20R%3D1%2B2(1%2B2%2B4%2B8)%3D31.%0A%20%20%20%20%24%24%0A%0A%20%20%20%20The%20final%20output%20can%20depend%20on%2031%20time%20positions%20while%20every%20kernel%20still%20has%20only%20three%20weights%20per%20input-output%20channel%20pair.%0A%0A%20%20%20%20%23%23%20Residual%20block%0A%0A%20%20%20%20A%20residual%20block%20returns%3A%0A%0A%20%20%20%20%24%24%0A%20%20%20%20%5Cmathbf%20y%3DF(%5Cmathbf%20x)%2B%5Cmathbf%20x.%0A%20%20%20%20%24%24%0A%0A%20%20%20%20%24F%24%20learns%20a%20correction%20instead%20of%20rebuilding%20the%20full%20representation.%20If%20channel%20counts%20differ%2C%20a%20learned%20projection%20aligns%20the%20shapes%20before%20addition.%0A%0A%20%20%20%20For%20forecasting%2C%20verify%20the%20receptive%20field%20against%20the%20longest%20plausible%20dependency.%20A%20network%20cannot%20learn%20from%20history%20that%20its%20connectivity%20never%20exposes.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20%23%23%20A%20practical%20example%0A%0A%20%20%20%20For%20machine-vibration%20monitoring%2C%20early%20filters%20can%20detect%20short%20oscillations%20while%20dilated%20layers%20connect%20those%20motifs%20across%20longer%20operating%20cycles.%20The%20architecture%20can%20classify%20a%20window%2C%20predict%20future%20readings%2C%20or%20flag%20unusual%20patterns.%0A%0A%20%20%20%20%23%23%20When%20to%20use%20it%0A%0A%20%20%20%20-%20Local%20motifs%20at%20multiple%20temporal%20scales%20are%20plausible.%0A%20%20%20%20-%20Predictions%20must%20be%20causal.%0A%20%20%20%20-%20Parallel%20training%20matters.%0A%20%20%20%20-%20You%20want%20explicit%20control%20over%20maximum%20context.%0A%20%20%20%20-%20Sequences%20are%20long%20enough%20to%20challenge%20simple%20dense%20models%20but%20attention%20is%20unnecessary%20or%20costly.%0A%0A%20%20%20%20%23%23%20When%20to%20avoid%20it%0A%0A%20%20%20%20-%20Relevant%20context%20is%20longer%20than%20the%20designed%20receptive%20field.%0A%20%20%20%20-%20Irregular%20event%20timing%20is%20not%20represented%20properly.%0A%20%20%20%20-%20Direct%20content-based%20interaction%20between%20arbitrary%20positions%20is%20central.%0A%20%20%20%20-%20Data%20is%20too%20small%20to%20justify%20a%20learned%20sequence%20architecture.%0A%20%20%20%20-%20A%20naive%20seasonal%20or%20classical%20model%20already%20solves%20the%20forecasting%20task.%0A%0A%20%20%20%20%23%23%20Data%20preparation%0A%0A%20%20%20%20Split%20by%20time%20or%20entity%20before%20creating%20overlapping%20windows.%20Fit%20scaling%20only%20on%20training%20periods.%20Include%20masks%20or%20elapsed-time%20features%20for%20irregular%20sampling.%20Confirm%20that%20every%20feature%20in%20a%20window%20would%20exist%20at%20prediction%20time.%0A%0A%20%20%20%20%23%23%20Important%20controls%0A%0A%20%20%20%20%7C%20Control%20%7C%20Role%20%7C%0A%20%20%20%20%7C---------%7C------%7C%0A%20%20%20%20%7C%20Kernel%20size%20%7C%20Local%20pattern%20width%20%7C%0A%20%20%20%20%7C%20Dilation%20schedule%20%7C%20How%20quickly%20history%20expands%20%7C%0A%20%20%20%20%7C%20Number%20of%20blocks%2Fchannels%20%7C%20Capacity%20%7C%0A%20%20%20%20%7C%20Dropout%20and%20weight%20decay%20%7C%20Regularization%20%7C%0A%20%20%20%20%7C%20Receptive%20field%20%7C%20Must%20cover%20plausible%20dependencies%20%7C%0A%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20Notebook%20%E2%80%94%20seeing%20the%20receptive%20field%0A%0A%20%20%20%20**Question%3A**%20how%20do%20causal%20dilation%20and%20depth%20decide%20which%20history%20can%20affect%20a%20prediction%3F%0A%0A%20%20%20%20This%20notebook%20isolates%20the%20architecture's%20connectivity.%20It%20does%20not%20train%20a%20large%20network.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_()%3A%0A%20%20%20%20import%20numpy%20as%20np%0A%20%20%20%20import%20matplotlib.pyplot%20as%20plt%0A%0A%20%20%20%20def%20receptive_positions(output_position%2C%20kernel_size%2C%20dilations)%3A%0A%20%20%20%20%20%20%20%20_positions%20%3D%20%7Boutput_position%7D%0A%20%20%20%20%20%20%20%20history%20%3D%20%5B_positions%5D%0A%20%20%20%20%20%20%20%20for%20dilation%20in%20reversed(dilations)%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20_positions%20%3D%20%7B%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20position%20-%20offset%20*%20dilation%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20for%20position%20in%20_positions%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20for%20offset%20in%20range(kernel_size)%0A%20%20%20%20%20%20%20%20%20%20%20%20%7D%0A%20%20%20%20%20%20%20%20%20%20%20%20history.append(_positions)%0A%20%20%20%20%20%20%20%20return%20list(reversed(history))%0A%0A%20%20%20%20return%20plt%2C%20receptive_positions%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20The%20setup%20traces%20which%20original%20time%20positions%20can%20reach%20one%20output%20through%20four%0A%20%20%20%20dilated%20layers.%20We%20first%20draw%20these%20dependencies%20as%20a%20graph%3B%20every%20edge%20points%20from%0A%20%20%20%20past%20information%20toward%20the%20current%20prediction%2C%20preserving%20causality.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(plt%2C%20receptive_positions)%3A%0A%20%20%20%20kernel_size%20%3D%203%0A%20%20%20%20dilations%20%3D%20%5B1%2C%202%2C%204%2C%208%5D%0A%20%20%20%20history%20%3D%20receptive_positions(output_position%3D30%2C%20kernel_size%3Dkernel_size%2C%20dilations%3Ddilations)%0A%20%20%20%20fig%2C%20ax%20%3D%20plt.subplots(figsize%3D(11%2C%204))%0A%20%20%20%20for%20layer%2C%20_positions%20in%20enumerate(history)%3A%0A%20%20%20%20%20%20%20%20ax.scatter(sorted(_positions)%2C%20%5Blayer%5D%20*%20len(_positions)%2C%20s%3D35)%0A%20%20%20%20ax.set(%0A%20%20%20%20%20%20%20%20yticks%3Drange(len(history))%2C%0A%20%20%20%20%20%20%20%20yticklabels%3D%5Bf%22layer%20%7Bi%7D%22%20for%20i%20in%20range(len(history))%5D%2C%0A%20%20%20%20)%0A%20%20%20%20ax.set(xlabel%3D%22sequence%20position%22%2C%20title%3D%22Positions%20connected%20to%20output%20at%20t%3D30%22)%0A%20%20%20%20ax.invert_yaxis()%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20The%20graph%20makes%20the%20expanding%20history%20visible.%20The%20final%20calculation%20lists%20the%0A%20%20%20%20exact%20positions%20and%20counts%20them%2C%20checking%20the%20receptive-field%20formula%20against%20the%0A%20%20%20%20dependency%20trace%20rather%20than%20accepting%20the%20formula%20on%20faith.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0A%40app.cell%0Adef%20_(receptive_positions)%3A%0A%20%20%20%20for%20schedule%20in%20(%5B1%2C%201%2C%201%2C%201%5D%2C%20%5B1%2C%202%2C%204%2C%208%5D%2C%20%5B1%2C%202%2C%204%2C%208%2C%2016%5D)%3A%0A%20%20%20%20%20%20%20%20_positions%20%3D%20receptive_positions(100%2C%20kernel_size%3D3%2C%20dilations%3Dschedule)%5B0%5D%0A%20%20%20%20%20%20%20%20print(%0A%20%20%20%20%20%20%20%20%20%20%20%20f%22dilations%3D%7Bschedule%7D%3A%20receptive%20field%20spans%20%22%0A%20%20%20%20%20%20%20%20%20%20%20%20f%22%7Bmax(_positions)%20-%20min(_positions)%20%2B%201%7D%20steps%22%0A%20%20%20%20%20%20%20%20)%0A%20%20%20%20return%0A%0A%0A%40app.cell(hide_code%3DTrue)%0Adef%20_(mo)%3A%0A%20%20%20%20mo.md(r%22%22%22%0A%20%20%20%20---%0A%0A%20%20%20%20%23%23%20Evaluation%20and%20diagnosis%0A%0A%20%20%20%20Compare%20against%20last-value%2C%20seasonal%2C%20linear%2C%20and%20tree-based%20baselines%20where%20appropriate.%20Evaluate%20across%20multiple%20future%20windows%20and%20regimes.%20Ablate%20context%20length%3A%20if%20performance%20does%20not%20change%2C%20the%20large%20receptive%20field%20may%20not%20be%20useful.%0A%0A%20%20%20%20Check%20boundary%20padding%2C%20latency%20for%20streaming%20use%2C%20and%20whether%20predictions%20shift%20when%20unavailable%20future%20values%20are%20deliberately%20perturbed.%20That%20last%20test%20can%20expose%20leakage.%0A%0A%20%20%20%20%23%23%20Cost%20profile%0A%0A%20%20%20%20Training%20parallelizes%20across%20sequence%20positions%20better%20than%20recurrent%20networks.%20Inference%20can%20be%20efficient%2C%20especially%20with%20caching%2C%20but%20naive%20recomputation%20over%20long%20windows%20wastes%20work.%0A%0A%20%20%20%20%23%23%20Related%20models%0A%0A%20%20%20%20-%20CNNs%20share%20convolutional%20filters%20but%20usually%20target%20spatial%20rather%20than%20temporal%20structure.%0A%20%20%20%20-%20RNN%2C%20LSTM%2C%20and%20GRU%20compress%20history%20into%20recurrent%20state.%0A%20%20%20%20-%20**Transformer**%20uses%20attention%20for%20flexible%20position-to-position%20interaction.%0A%20%20%20%20-%20State-space%20models%20offer%20another%20route%20to%20long%20efficient%20context.%0A%0A%20%20%20%20%23%23%20What%20the%20receptive%20field%20controls%0A%0A%20%20%20%20Choose%20a%20TCN%20when%20%22patterns%20across%20known%20temporal%20scales%22%20is%20a%20better%20description%20of%20the%20problem%20than%20%22every%20position%20may%20need%20to%20look%20directly%20at%20every%20other%20position.%22%0A%0A%20%20%20%20%23%23%20Concept%20references%0A%0A%20%20%20%20-%20%5BData%20Leakage%5D(%2Fconcepts%2Fdata_leakage)%20%E2%80%94%20causal%20windows%20must%20not%20contain%20future%20information.%0A%20%20%20%20-%20%5BBackpropagation%5D(%2Fconcepts%2Fbackpropagation)%20%E2%80%94%20how%20convolution%20filters%20receive%20gradients.%0A%20%20%20%20%22%22%22)%0A%20%20%20%20return%0A%0A%0Aif%20__name__%20%3D%3D%20%22__main__%22%3A%0A%20%20%20%20app.run()%0A
c8c62069edf51e50919a258a5bf964a5