Init loss $=\log C$ ($\ln10=2.30$). Example (3.2, 5.1, −1.7): $p=(0.13,0.87,0.00)$, $L=2.04$.
Cat example $=2.9$; init $=C-1$. Zero once margins hold (stops learning); softmax never stops. $2W$ keeps $L=0$ → need $\lambda R(W)$.
$\beta_1$ 0.9, $\beta_2$ 0.999, start lr $10^{-3}$ or $5\cdot10^{-4}$. Numerical gradient check: $(1.25322-1.25347)/10^{-4}=-2.5$.
downstream = local × upstream. add distributes, mul swaps, max routes. $q=x+y,\ f=qz$: $\partial f/\partial x=z=-4$.
$$y=xW:\quad \frac{\partial L}{\partial x}=\frac{\partial L}{\partial y}W^T,\qquad \frac{\partial L}{\partial W}=x^T\frac{\partial L}{\partial y}$$Sigmoid: $\sigma'=\sigma(1-\sigma)=0.73\cdot0.27=0.20$; $dw=[-0.2,-0.39,0.2]$, $dx=[0.39,-0.59]$.
3×32×32, 10×(5×5), S1 P2 → 10×32×32, 760 params, 768,000 MACs. "Same" $P=(K-1)/2$. Receptive field $1+L(K-1)$. VGG: three 3×3 = one 7×7, $27C^2$ vs $49C^2$. ResNet $H(x)=F(x)+x$.
TP: class ok and IoU ≥ τ. Duplicate = FP. No TN. Course table: Cat 0.6667, Dog 0.5, Bicycle 0.3333 → mAP 0.5000. Bicycle IoU 0.74 fails τ 0.75.
$W$: $4h\times(h+d)$. $\partial c_t/\partial c_{t-1}=\mathrm{diag}(f)$ → "uninterrupted flow, like ResNet". Clip for exploding.
$\sqrt D$: $\mathrm{Var}(q\cdot k)=D$ → avoid softmax saturation. Permutation-equivariant → positional encoding. Mask future with $-\infty$. Block = MHSA → +res → LN → MLP(D→4D→D) → +res → LN; 6 matmuls; $O(N^2)$. ViT: $N=HW/P^2$ patches (224/16 → 196 tokens of 768), CLS token, low inductive bias → needs big data.
"Training minimises a loss by gradient descent; backprop is just the chain rule on the computational graph. CNNs share small filters across positions; LSTMs keep a cell state whose gradient passes through an element-wise gate; attention lets every token read every other token in one step at $O(N^2)$ cost." Overfitting = low train / high test error; underfitting = both high and "cannot be fixed by more epochs".
$e=SP-PV$. P only → steady-state error (0 output at $e=0$); I removes it; D damps, amplifies noise. Tune $K_p$, then $K_d$, then $K_i$ (0.0001). $K_p=125/3500=0.0357$. Example (2.0,0.5,0.1), $\Delta t$ 0.1, $e=-0.48$, $e_{prev}=-0.30$, $\sum=-0.12$: $-0.96-0.06-0.18=\mathbf{-1.20}$.
Normal form: vertical lines finite. (3,3),(4,3),(5,3) → $A(3,90°)=3$. RANSAC $P$ 0.99, $p$ 0.5: $k$=2→17, 3→35, 4→72. Reject lane parabola $|a|\ge0.003$.
(700,400), $Z$ 10, $f$ 800, $c$ (640,360) → (0.75, 0.5, 10). Stereo $d=80$, $f_x$ 795, $B$ 0.2 → $Z=1.9875$ m. One image: 2 eq, 3 unknowns. Quality = reprojection error.
Output $S\times S\times(5B+C)$: 7×7×30. Loss: $\sqrt w,\sqrt h$; $\lambda_{noobj}=0.5$; NMS IoU > 0.5. MIO: in lane if $x_L(y)\le x\le x_R(y)$, $x(y)=(y-b)/m$; MIO = argmax $y_{bottom}$. Three-car example: Car 3 (x 500) out of lane → Car 2 (290 > 260). FCW: tracks (confirm [2 3], delete 5), closest in lane; $d=1.2v+v^2/(2\cdot0.4\cdot9.8)$ → 24.8 m at 10 m/s.
Derivation: $N(\mu_p,p)\times N(z,r)$ → $\frac1{\sigma^2}=\frac1p+\frac1r$, $\mu=\frac{r\mu_p+pz}{p+r}$; set $K=\frac{p}{p+r}$. $K\to1$ trust sensor, $K\to0$ trust prediction; $0 Fusion: $z_f=\frac{\sum z_i/r_i}{\sum 1/r_i},\ r_f=\frac1{\sum1/r_i}$; (0.9,1.1),(1,4) → 0.94, 0.8; prior 10.1 → $K=0.927$, $x=0.87$, $p=0.74$. $r_i$ never changes during filtering.
"Bang-bang chooses a direction; PID chooses how much." "Hough votes in $(\rho,\theta)$ because slope is infinite for vertical lines." "RANSAC keeps the model most points agree with; least squares is pulled by outliers." "The Kalman filter is a recursive Bayesian estimator: predict widens, update shrinks." "MIO comes from confirmed tracks because detections flicker."
Pull-up button pressed = LOW. Never delay(), use millis(). Pooling has no parameters; PID I-term needs anti-windup in practice. Behaviour cloning is regression (ELU + regression layer), not softmax. Gazebo = world, RViz = belief.
Error ≤ $s/2$, variance $s^2/12$. Weights per-channel, activations per-tensor, bias int32 with $s_b=s_as_x$. PTQ (observe min/max) vs QAT (fake-quant nodes). MCU: int32 accumulator, shift >> 7.
w = !(~X*X)*~X*y (~ transpose, ! Gauss–Jordan inverse, pivot < 1e-6 → abort). (1,2),(2,2.5),(3,3.5): $m=4.5/6=0.75$, $c=1.17$. Quadratic (1,2),(2,3),(3,5): $\beta=[2,-0.5,0.5]$. $\begin{bmatrix}2&1\\5&3\end{bmatrix}^{-1}=\begin{bmatrix}3&-1\\-5&2\end{bmatrix}$. "Linear in parameters, nonlinear in features."
1000 mAh, 50 mAh every 2 h → 1.67 d. 1200 mAh, 40 mAh, 5 d → every 4 h. Uno 500 mAh: awake 50 mA → 10 h; asleep 0.1 mA → 30 days; inference 45 mA·0.8 s = 0.01 mAh. $\lambda=1/12$: $S(2)$ 0.846, $S(8.3)$ 0.5, $S(24)$ 0.135; run if $S<0.3$ (t > 14.4 h). Adaptive interval $60TE/B$: 180 s @5000, 900 s @1000. Morning 7 min + night 22.5 min → 67 events → 2010 mAh.
Raspberry Pi: Interpreter → allocate_tensors → set_tensor → invoke → get_tensor; detection input [1,320,320,3] uint8, outputs boxes/classes/scores, thr 0.3, EfficientDet-Lite0. Arduino only if model < 20 KB. Hierarchy: Arduino wakes Pi at 10–30 cm.
| UART | I²C | SPI | |
|---|---|---|---|
| wires | 3 (Rx,Tx,GND) | 2 (SDA,SCL) | 4 + n (SCK,MOSI,MISO,SS) |
| clock | none (baud) | master | master |
| duplex | full | half | full |
| speed | 9600/115200 | 100k/400k/3.4M | ~10 MHz |
| notes | RS-232 ±3–25 V, MAX232 | 7-bit → 128 addr; START SDA↓ while SCL high; ACK SDA low; 4.7 kΩ pull-ups | SS low selects; byte per 8 clocks |
Ride OFF→ON 50%→OFF after 20 s. Fan Idle→30%→60%→100%. Traffic Green 120 → Yellow 30 → 3 blinks (6 states) → Red 120; pedestrian PB only in Green with > 30 s left. Python: Enum + loop; (value+1) % 4.
"Quantisation stores each weight as an integer plus a shared scale and zero-point; integer inference needs only an int32 accumulator and a rescale." "Idle current dominates the battery: sleep and wake on interrupts." "$S(t)$ is a survival probability, not a density." "Least squares has a closed form, $(X^TX)^{-1}X^Ty$, so a microcontroller can learn without gradient descent."
DNA = text in A,C,G,T; variant = changed letter (SNV) or small insert/delete (indel); monogenic = one gene (CF, sickle cell, PKU, MPS I, Duchenne, FMF, β-thal, Alport, Bardet–Biedl); motif = short meaningful pattern (splice site, TF binding site); reverse complement = reverse string, swap A↔T, C↔G (AACG → CGTT); ClinVar/ClinGen = curated databases; ACMG/AMP = 5-tier interpretation guideline; gnomAD = healthy-population variants.
Conv1D filters = motif detectors (local, position-invariant); max pooling = "motif present somewhere"; BiLSTM = context both upstream and downstream; combination claims local + long-range + contextual dependencies. Class imbalance: ×500 augmentation + class-inverse weights ($w_c\propto1/n_c$) + stratified batches; evaluate with weighted F1 and AUC-PR (PR more informative than ROC when positives are rare).
"Single-gene diseases are diagnosed by finding which DNA spelling change is harmful; yield is 25–50%. The authors cut a 101-letter window around known variants, place it in random background 500 times with augmentation, and train a 1-D CNN (motifs) plus a bidirectional LSTM (context) with class-weighted cross-entropy and Adam. It reaches 94.7% accuracy, F1 0.93, AUC-PR 0.98 on their synthetic test set; MPS I and PKU are weakest. The honest limit: synthetic negatives and no external validation, so clinical performance is unknown."
$\partial\mathrm{KL}/\partial z_s=(p_s-p_t)/T$. $T\to\infty$ uniform; $T\to0$ one-hot. Softmax over feature dims is a heuristic; AS-2 justifies it.
Tactile sensors = camera inside a soft gel; optics, gel, light differ per device → same material, different images → models do not transfer. Words ("rough, soft, slippery") are sensor-agnostic → use a frozen language model as teacher; distil its embedding into a ViT; afterwards train only a tiny head per dataset/sensor. Language only at training time.
AS-1: richer seq2seq teacher (BART) gives richer supervision. AS-2: feature-level KD ≫ cosine (direction only, flat near alignment) ≫ DKD (logits need a shared classifier; modalities differ). AS-3: moderate α, T; too supervised (α 0.8 → 80.68) or extreme T worse. AS-4: small batch regularises; lr 3e-5 unstable, 1e-5 slow. UMAP: tighter clusters. Grad-CAM: texture/edges, specular regions.
"Robots need touch; tactile cameras differ, so models fail across sensors. The authors use language as a sensor-agnostic teacher: a frozen BART embeds touch descriptions, a ViT is trained to match them with a temperature-softened KL loss (T 3.5) mixed 0.25/0.75 with cross-entropy, then frozen; only a tiny head is trained per task. On 39K relabelled DIGIT samples (32 classes) it reaches 95% at 100 shots, +13.3% average cross-sensor gain to GelSight, 98.8% on HCT, and beats vision-only using touch alone."
Two nonstationarities: inherent (other agents learn; solved by CTDE critic) vs system (world changes via μ; solved by recurrent policy + DR). Proposition 1: under A1 (others = environment) and A2 (observation bijective), an RNN policy over history can infer $P_\mu$ hence μ (Glivenko–Cantelli), so recurrent CTDE MARL solves the game. Gaps: i.i.d. assumption, 30 steps, no learning guarantee.
RMRand > MRand (memory helps under DR). RARand ≫ RCRand, MRand (memory in the actor is what matters; ≈ RMRand). MRand > MADDPG (DR helps even without memory). Sim: similar completion, RMRand ≈ RARand fastest. Real world: RMRand best completion and time; MADDPG without DR fails completely. Noise intuition: σ/√n (3 m → 1.34 m over 5 steps).
"Two drones carry goods on ropes to two markers. Training in reality is unsafe, so they train in AirSim with domain randomization over six parameters so the real world is one sample of the simulated distribution. Policies are R-MADDPG: LSTM memory in actor and critic so hidden conditions can be inferred from history; perception is a separate CNN detector. The memory-based randomised policy transfers directly to real F450/PX4 drones with the best completion rate; the actor's memory matters more than the critic's; without randomization the policy fails."