Colour cheat sheet · 6 modules, blocks stacked one below the other · A4. Print with "Background graphics" enabled.
1

Deep Learning foundations

MAI590/DEN790 · losses, optimisers, backprop, CNN arithmetic, detection metrics, LSTM, attention
formulanumbers to quotetrap / critiquesay this in the exam

Softmax & cross-entropy

$$p_k=\frac{e^{s_k}}{\sum_j e^{s_j}},\qquad L=-\log p_y,\qquad \frac{\partial L}{\partial s_j}=p_j-\mathbb 1[j=y]$$

Init loss $=\log C$ ($\ln10=2.30$). Example (3.2, 5.1, −1.7): $p=(0.13,0.87,0.00)$, $L=2.04$.

Multiclass SVM (hinge)

$$L_i=\sum_{j\neq y}\max(0,\,s_j-s_y+1)$$

Cat example $=2.9$; init $=C-1$. Zero once margins hold (stops learning); softmax never stops. $2W$ keeps $L=0$ → need $\lambda R(W)$.

Optimisers

$$\text{SGD: }w\leftarrow w-\alpha\nabla L\qquad\text{Momentum: }v\leftarrow\rho v+\nabla L,\ w\leftarrow w-\alpha v$$ $$\text{Adam: }m\leftarrow\beta_1m+(1-\beta_1)g,\ v\leftarrow\beta_2v+(1-\beta_2)g^2,\ \hat m=\tfrac{m}{1-\beta_1^t},\ \hat v=\tfrac{v}{1-\beta_2^t},\ w\leftarrow w-\alpha\tfrac{\hat m}{\sqrt{\hat v}+\epsilon}$$

$\beta_1$ 0.9, $\beta_2$ 0.999, start lr $10^{-3}$ or $5\cdot10^{-4}$. Numerical gradient check: $(1.25322-1.25347)/10^{-4}=-2.5$.

Backprop rules

downstream = local × upstream. add distributes, mul swaps, max routes. $q=x+y,\ f=qz$: $\partial f/\partial x=z=-4$.

$$y=xW:\quad \frac{\partial L}{\partial x}=\frac{\partial L}{\partial y}W^T,\qquad \frac{\partial L}{\partial W}=x^T\frac{\partial L}{\partial y}$$

Sigmoid: $\sigma'=\sigma(1-\sigma)=0.73\cdot0.27=0.20$; $dw=[-0.2,-0.39,0.2]$, $dx=[0.39,-0.59]$.

Convolution arithmetic

$$W'=\Big\lfloor\frac{W-K+2P}{S}\Big\rfloor+1,\qquad \#\text{params}=C_{out}(C_{in}K^2+1),\qquad \text{MACs}=C_{in}K^2\cdot W'H'C_{out}$$

3×32×32, 10×(5×5), S1 P2 → 10×32×32, 760 params, 768,000 MACs. "Same" $P=(K-1)/2$. Receptive field $1+L(K-1)$. VGG: three 3×3 = one 7×7, $27C^2$ vs $49C^2$. ResNet $H(x)=F(x)+x$.

Numbers to quote

ILSVRC top-5AlexNet 8L · VGG 7.3% · ResNet-152 3.57%CIFAR-1050k/10k, 32×32×3 = 3072Max-pool 2×2/2[[1,1,2,4],[5,6,7,8],[3,2,1,0],[1,2,3,4]] → [[6,8],[3,4]]Dropoutp 0.5; inverted: ÷(1−p) in trainingConv 7×7 backward∂L/∂b = 8; ∂L/∂W = [[20,10,2],[5,18,16],[15,10,4]]Transformer sizes12L/213M · GPT-2 48L/1.5B · GPT-3 96L/175B

Detection metrics

$$\mathrm{IoU}=\frac{|P\cap G|}{|P\cup G|},\quad P=\frac{TP}{TP+FP},\quad R=\frac{TP}{TP+FN},\quad AP=\sum_i\Delta R_iP_i,\quad \mathrm{mAP}=\tfrac1C\sum AP_c$$

TP: class ok and IoU ≥ τ. Duplicate = FP. No TN. Course table: Cat 0.6667, Dog 0.5, Bicycle 0.3333 → mAP 0.5000. Bicycle IoU 0.74 fails τ 0.75.

RNN → LSTM

$$h_t=\tanh(W_{hh}h_{t-1}+W_{xh}x_t)\quad\Rightarrow\quad \prod_t \mathrm{diag}(\tanh')W^T\ \text{vanishes/explodes}$$ $$\begin{pmatrix}i\\f\\o\\g\end{pmatrix}=\begin{pmatrix}\sigma\\\sigma\\\sigma\\\tanh\end{pmatrix}W\begin{pmatrix}h_{t-1}\\x_t\end{pmatrix},\ c_t=f\odot c_{t-1}+i\odot g,\ h_t=o\odot\tanh c_t$$

$W$: $4h\times(h+d)$. $\partial c_t/\partial c_{t-1}=\mathrm{diag}(f)$ → "uninterrupted flow, like ResNet". Clip for exploding.

Attention & ViT

$$Q=XW_Q,\ K=XW_K,\ V=XW_V,\qquad Y=\mathrm{softmax}\!\Big(\frac{QK^T}{\sqrt D}\Big)V$$

$\sqrt D$: $\mathrm{Var}(q\cdot k)=D$ → avoid softmax saturation. Permutation-equivariant → positional encoding. Mask future with $-\infty$. Block = MHSA → +res → LN → MLP(D→4D→D) → +res → LN; 6 matmuls; $O(N^2)$. ViT: $N=HW/P^2$ patches (224/16 → 196 tokens of 768), CLS token, low inductive bias → needs big data.

Say this in the exam

"Training minimises a loss by gradient descent; backprop is just the chain rule on the computational graph. CNNs share small filters across positions; LSTMs keep a cell state whose gradient passes through an element-wise gate; attention lets every token read every other token in one step at $O(N^2)$ cost." Overfitting = low train / high test error; underfitting = both high and "cannot be fixed by more epochs".

Viva cheat sheet · Deep Learningpage 1 of 6
2

Intelligent Robots foundations

MAI675/DEN775 · PID, Hough, RANSAC, camera, YOLO/MIO, Kalman, ROS 2

The loop and the pipeline

Sense→Compute→Actuate⟳image→grey→edges (Sobel/Canny)→ROI→Hough→filter L/R→centre, error→PID→servo / wheels

PID

$$u=K_pe+K_i\!\int\! e\,dt+K_d\frac{de}{dt}\qquad u[n]=K_pe[n]+K_i\sum e[i]\Delta t+K_d\frac{e[n]-e[n-1]}{\Delta t}$$

$e=SP-PV$. P only → steady-state error (0 output at $e=0$); I removes it; D damps, amplifies noise. Tune $K_p$, then $K_d$, then $K_i$ (0.0001). $K_p=125/3500=0.0357$. Example (2.0,0.5,0.1), $\Delta t$ 0.1, $e=-0.48$, $e_{prev}=-0.30$, $\sum=-0.12$: $-0.96-0.06-0.18=\mathbf{-1.20}$.

Hough & RANSAC

$$\rho=x\cos\theta+y\sin\theta,\quad m=-\cot\theta,\ b=\rho/\sin\theta\qquad S=\frac{\log(1-P)}{\log(1-p^k)}$$

Normal form: vertical lines finite. (3,3),(4,3),(5,3) → $A(3,90°)=3$. RANSAC $P$ 0.99, $p$ 0.5: $k$=2→17, 3→35, 4→72. Reject lane parabola $|a|\ge0.003$.

Camera & stereo

$$\lambda\begin{bmatrix}u\\v\\1\end{bmatrix}=K[R|t]\begin{bmatrix}X\\Y\\Z\\1\end{bmatrix},\ K=\begin{bmatrix}f_x&s&c_x\\0&f_y&c_y\\0&0&1\end{bmatrix},\ X=\frac{(u-c_x)Z}{f_x},\ Z=\frac{f_xB}{d}$$

(700,400), $Z$ 10, $f$ 800, $c$ (640,360) → (0.75, 0.5, 10). Stereo $d=80$, $f_x$ 795, $B$ 0.2 → $Z=1.9875$ m. One image: 2 eq, 3 unknowns. Quality = reprojection error.

YOLO & MIO

Output $S\times S\times(5B+C)$: 7×7×30. Loss: $\sqrt w,\sqrt h$; $\lambda_{noobj}=0.5$; NMS IoU > 0.5. MIO: in lane if $x_L(y)\le x\le x_R(y)$, $x(y)=(y-b)/m$; MIO = argmax $y_{bottom}$. Three-car example: Car 3 (x 500) out of lane → Car 2 (290 > 260). FCW: tracks (confirm [2 3], delete 5), closest in lane; $d=1.2v+v^2/(2\cdot0.4\cdot9.8)$ → 24.8 m at 10 m/s.

Kalman filter · five scalar equations and their origin

$$\underbrace{\mu_p=\mu+v\Delta t+\tfrac12a\Delta t^2}_{\text{mechanics}}\quad \underbrace{p_p=p+q}_{\text{variances add}}\quad \underbrace{K=\frac{p_p}{p_p+r}}_{\text{confidence}}\quad \underbrace{\mu=\mu_p+K(z-\mu_p)}_{\text{Bayes mean}}\quad \underbrace{p=(1-K)p_p}_{\text{Bayes variance}}$$

Derivation: $N(\mu_p,p)\times N(z,r)$ → $\frac1{\sigma^2}=\frac1p+\frac1r$, $\mu=\frac{r\mu_p+pz}{p+r}$; set $K=\frac{p}{p+r}$. $K\to1$ trust sensor, $K\to0$ trust prediction; $0

Fusion: $z_f=\frac{\sum z_i/r_i}{\sum 1/r_i},\ r_f=\frac1{\sum1/r_i}$; (0.9,1.1),(1,4) → 0.94, 0.8; prior 10.1 → $K=0.927$, $x=0.87$, $p=0.74$. $r_i$ never changes during filtering.

Numbers to quote

Ultrasonicd = t·0.034/2 cm (2000 µs → 34 cm)L293DIN 10 fwd, 01 rev, 00 stop; EN = PWMLine bar I²Cbar 0x3E, robot 81, request 240Grey0.299R+0.587G+0.114B; (120,200,80) → 162Sobel ex.Gx 275, Gy 145 → M 311, θ 27.8°HoughLinesPrho 2, θ π/180, thr 50, minLen 10, gap 30BC net24/36/48 (5×5 s2), 64/64, FC 1254/1254/256/1, lr 1e-4ROS 2Jazzy · colcon · /thing_on Bool q10 · frame "map"

Say this in the exam

"Bang-bang chooses a direction; PID chooses how much." "Hough votes in $(\rho,\theta)$ because slope is infinite for vertical lines." "RANSAC keeps the model most points agree with; least squares is pulled by outliers." "The Kalman filter is a recursive Bayesian estimator: predict widens, update shrinks." "MIO comes from confirmed tracks because detections flicker."

Traps

Pull-up button pressed = LOW. Never delay(), use millis(). Pooling has no parameters; PID I-term needs anti-windup in practice. Behaviour cloning is regression (ELU + regression layer), not softmax. Gazebo = world, RViz = belief.

Viva cheat sheet · Intelligent Robotspage 2 of 6
3

Edge AI foundations

AIRE325 / MAI633 · quantisation, regression on MCUs, energy, TinyML, protocols, state charts

Quantisation

$$x_q=\mathrm{round}\!\Big(\frac{x}{s}\Big)+z,\qquad x\approx s(x_q-z),\qquad s=\frac{x_{max}-x_{min}}{255}$$ $$y_q=\frac{s_as_x}{s_y}(a_q-z_a)(x_q-z_x)+\frac{s_b}{s_y}b_q+z_y$$

Error ≤ $s/2$, variance $s^2/12$. Weights per-channel, activations per-tensor, bias int32 with $s_b=s_as_x$. PTQ (observe min/max) vs QAT (fake-quant nodes). MCU: int32 accumulator, shift >> 7.

Numbers to quote

Slide examplex 4.0 → x_q 102; a_q 64; b_q 13; y_q 112.8 → 11.05 (float 11.0)Exercises3.2/0.04 → 80 · 0.05(100−128) = −1.4 · 6.375/255 = 0.025Accuracy dropMobileNet 71.9→71.0 · KWS 94.0→93.8 · CIFAR 82.4→82.3Uno2 KB SRAM, 32 KB flash; int8+PROGMEM ≈ 4× neurons (50→200)Nano 33 BLE Sense256 KB / 1 MB; IMU, mic, T/H, pressureGesture lab119×6 samples, thr 2.5 g, 50-15-2, 600 ep, arena 8 KBKWS lab2500 ms @ 16 kHz, MFCC, Edge Impulse

Regression on a microcontroller

$$\hat\beta=(X^TX)^{-1}X^Ty\qquad m=\frac{n\sum xy-\sum x\sum y}{n\sum x^2-(\sum x)^2},\ c=\bar y-m\bar x$$

w = !(~X*X)*~X*y (~ transpose, ! Gauss–Jordan inverse, pivot < 1e-6 → abort). (1,2),(2,2.5),(3,3.5): $m=4.5/6=0.75$, $c=1.17$. Quadratic (1,2),(2,3),(3,5): $\beta=[2,-0.5,0.5]$. $\begin{bmatrix}2&1\\5&3\end{bmatrix}^{-1}=\begin{bmatrix}3&-1\\-5&2\end{bmatrix}$. "Linear in parameters, nonlinear in features."

Energy budgets

$$\text{uptime}=\frac{\text{capacity}}{\text{inf/day}\times\text{mAh/inf}},\qquad S(t)=e^{-\lambda t},\ \lambda=\tfrac1{\text{mean}},\qquad E[N]=\frac{T}{E[T]}$$

1000 mAh, 50 mAh every 2 h → 1.67 d. 1200 mAh, 40 mAh, 5 d → every 4 h. Uno 500 mAh: awake 50 mA → 10 h; asleep 0.1 mA → 30 days; inference 45 mA·0.8 s = 0.01 mAh. $\lambda=1/12$: $S(2)$ 0.846, $S(8.3)$ 0.5, $S(24)$ 0.135; run if $S<0.3$ (t > 14.4 h). Adaptive interval $60TE/B$: 180 s @5000, 900 s @1000. Morning 7 min + night 22.5 min → 67 events → 2010 mAh.

TinyML flow

train (PC)→prune · cluster · quantise→.tflite→xxd → model.h→TFLite Micro:ErrorReporterGetModel + versionOpResolver<N>tensor arenaInterpreterAllocateTensorsinput → Invoke → output

Raspberry Pi: Interpreter → allocate_tensors → set_tensor → invoke → get_tensor; detection input [1,320,320,3] uint8, outputs boxes/classes/scores, thr 0.3, EfficientDet-Lite0. Arduino only if model < 20 KB. Hierarchy: Arduino wakes Pi at 10–30 cm.

Protocols

UARTI²CSPI
wires3 (Rx,Tx,GND)2 (SDA,SCL)4 + n (SCK,MOSI,MISO,SS)
clocknone (baud)mastermaster
duplexfullhalffull
speed9600/115200100k/400k/3.4M~10 MHz
notesRS-232 ±3–25 V, MAX2327-bit → 128 addr; START SDA↓ while SCL high; ACK SDA low; 4.7 kΩ pull-upsSS low selects; byte per 8 clocks

State charts (method)

classify devices (in/out, dig/ana)→output string→count states→outputs/state→transitions (PB, After t, guard)

Ride OFF→ON 50%→OFF after 20 s. Fan Idle→30%→60%→100%. Traffic Green 120 → Yellow 30 → 3 blinks (6 states) → Red 120; pedestrian PB only in Green with > 30 s left. Python: Enum + loop; (value+1) % 4.

Say this in the exam

"Quantisation stores each weight as an integer plus a shared scale and zero-point; integer inference needs only an int32 accumulator and a rescale." "Idle current dominates the battery: sleep and wake on interrupts." "$S(t)$ is a survival probability, not a density." "Least squares has a closed form, $(X^TX)^{-1}X^Ty$, so a microcontroller can learn without gradient descent."

Viva cheat sheet · Edge AIpage 3 of 6
P1

CNN–BiLSTM for pathogenic variants

Abdelrehim & Mohamed 2026, Intelligent Systems with Applications 30:200654 · Liwa University

Pipeline

ClinVar / ClinGen variants→clean, title-case labels→101-bp window, random background ×500→reverse complement + k-mer jitter→+ random "Not a Disease"→A,C,G,T,N → 1,2,3,4,0 → one-hot→70/15/15 stratified, seed 42→multiscale Conv1D→max-pool→BiLSTM→attention/dense→softmax 21 classes

Equations

$$c_t^{(k)}=\sigma\Big(\sum_{i=0}^{f-1}w_{k,i}\cdot y_{t+i}+b_k\Big),\ t=1..L-f+1\ (1)\qquad p^{(k)}=\max_t c_t^{(k)}\ (2)$$ $$\overrightarrow d_t=\mathrm{LSTM}(\overrightarrow d_{t-1},y_t,\overrightarrow s_{t-1}),\ \overleftarrow d_t=\mathrm{LSTM}(\overleftarrow d_{t-1},y_t,\overleftarrow s_{t-1})\ (3,4)\qquad L=-\tfrac1N\sum_i\sum_c y_{i,c}\log\hat y_{i,c}\ (5)$$ $$\text{Spec}=\tfrac{TN}{TN+FP}\ (6)\quad \text{Acc}=\tfrac{TP+TN}{\text{all}}\ (7)\quad \text{Fallout}=\tfrac{FP}{FP+TN}\ (8)\quad LR^-=\tfrac{1-\text{Sens}}{\text{Spec}}\ (9)\quad NPV=\tfrac{TN}{TN+FN}\ (10)$$

Numbers to quote

Task20 monogenic diseases + "Not a Disease" = 21 classesWindow101 bp (±50), one-hot 101×4Augmentation×500 per mutation; RC + k-mer jitterTrainingAdam 1e-4 (β 0.9/0.999, ε 1e-8), batch 32, ≤50 ep, patience 7, best-val checkpoint, ≈12 ep, GPU, seed 42Resultsacc 94.7% · wF1 0.93±0.03 · AUC-PR 0.98±0.02 · ROC-AUC ≈1.00 · spec 0.98 · NPV >94%Hard classesMPS I (IDUA W402X), PKU (PAH R408W): acc 0.84, spec 0.889, 2 FP eachDiagnostic yield25–50% (exome 25–30%); ≈7,000 monogenic diseasesDeploymentFlask/Render web service "using LLMs"

Biology in one breath

DNA = text in A,C,G,T; variant = changed letter (SNV) or small insert/delete (indel); monogenic = one gene (CF, sickle cell, PKU, MPS I, Duchenne, FMF, β-thal, Alport, Bardet–Biedl); motif = short meaningful pattern (splice site, TF binding site); reverse complement = reverse string, swap A↔T, C↔G (AACG → CGTT); ClinVar/ClinGen = curated databases; ACMG/AMP = 5-tier interpretation guideline; gnomAD = healthy-population variants.

Why CNN + BiLSTM

Conv1D filters = motif detectors (local, position-invariant); max pooling = "motif present somewhere"; BiLSTM = context both upstream and downstream; combination claims local + long-range + contextual dependencies. Class imbalance: ×500 augmentation + class-inverse weights ($w_c\propto1/n_c$) + stratified batches; evaluate with weighted F1 and AUC-PR (PR more informative than ROC when positives are rare).

Critique (prepare calmly)

  • Negatives are random letters, positives are real motifs in random background → model may learn "real vs random". Authors admit it; fix: gnomAD/dbSNP benign variants.
  • 500 near-duplicate windows + random split → leakage between train/test; fix: split by mutation.
  • Ablation values printed as "[insert value]"; Table 2 rows have TP = 0 yet F1 0.93 (metrics from "hypothetical 19 negatives per class").
  • "Long-range" ≤ 50 bp; no benchmark vs CADD/SpliceAI; SHAP/PoSHAP and calibration are future work; LLM role unexplained.

Two-minute summary

"Single-gene diseases are diagnosed by finding which DNA spelling change is harmful; yield is 25–50%. The authors cut a 101-letter window around known variants, place it in random background 500 times with augmentation, and train a 1-D CNN (motifs) plus a bidirectional LSTM (context) with class-weighted cross-entropy and Adam. It reaches 94.7% accuracy, F1 0.93, AUC-PR 0.98 on their synthetic test set; MPS I and PKU are weakest. The honest limit: synthetic negatives and no external validation, so clinical performance is unknown."

Viva cheat sheet · Paper 1page 4 of 6
P2

Language-guided tactile representation learning

Mohsan, Ud Din, Xu, Abubakar, Hussain 2026 (arXiv 2609.14783) · KUCARS, Khalifa University

Pipeline

tactile image (DIGIT)→ViT student $f_\theta$ → $z_{tactile}$⇄ KLfrozen BART teacher → $z_{text}$+ CE on labels→$L=\alpha L_{CE}+(1-\alpha)L_{KD}$→freeze $\theta$→MLP head $h_\phi$ (0.017%)→material (touch only at test)

Equations

$$p_t^{(T)}=\mathrm{softmax}(z_{text}/T),\quad p_s^{(T)}=\mathrm{softmax}(z_{tactile}/T)\ (6)\qquad L_{KD\text{-}feat}=\mathrm{KL}(p_t^{(T)}\|p_s^{(T)})=\sum_j p_t\log\frac{p_t}{p_s}\ (7)$$ $$L_{student}=-\sum_c Y_T\log p_s(c)\ (8)\qquad L=\alpha L_{student}+(1-\alpha)L_{KD\text{-}feat}\ (9)\qquad \min_\phi \mathbb E[L_{CE}(h_\phi(f_\theta(x)),y)],\ \theta^*=\theta\ (10\text{–}12)$$

$\partial\mathrm{KL}/\partial z_s=(p_s-p_t)/T$. $T\to\infty$ uniform; $T\to0$ one-hot. Softmax over feature dims is a heuristic; AS-2 justifies it.

Numbers to quote

DatasetHCT 36K + SSVTP 4K = 39K DIGIT; 3 annotators; 32 classes: 12 distil / 20 unseenBest hyper-paramsα 0.25, T 3.5; fine-tune batch 32, lr 2e-5; RTX 4090, ≤100 epFew-shotK ∈ {0,10,100,1000}; ≈95% at 100-shotCross-sensorDIGIT → GelSight Hex: TAG +22% (0-shot), +18% (100), +9% (1000); ICRA18 +10.25%; avg +13.3%Table II (ours)TAG 61.20 · YCB 97.71 · ICRA18 73.86 · FEEL 86.88 · SSVTP 83.52 · HCT 98.80 · own 95.06Modalitieslang 60.38 · tactile 64.99 · vision 90.89 · ours 95.06AS-1 teacherBART 58.64 > DistilBERT 53.86 > RoBERTa 42.69AS-2 lossKD-feat 82.66 ≫ CosMin 58.64 ≫ DKD 34.02AS-4B 32 90.91 (512 → 45.44); lr 2e-5 95.06

Idea in one breath

Tactile sensors = camera inside a soft gel; optics, gel, light differ per device → same material, different images → models do not transfer. Words ("rough, soft, slippery") are sensor-agnostic → use a frozen language model as teacher; distil its embedding into a ViT; afterwards train only a tiny head per dataset/sensor. Language only at training time.

Ablation logic

AS-1: richer seq2seq teacher (BART) gives richer supervision. AS-2: feature-level KD ≫ cosine (direction only, flat near alignment) ≫ DKD (logits need a shared classifier; modalities differ). AS-3: moderate α, T; too supervised (α 0.8 → 80.68) or extreme T worse. AS-4: small batch regularises; lr 3e-5 unstable, 1e-5 slow. UMAP: tighter clusters. Grad-CAM: texture/edges, specular regions.

Critique (prepare calmly)

  • Distilled on one sensor; cross-sensor test classes overlap training classes (authors say so). Novel sensor + novel material untested.
  • Baselines train only on target data; proposed model gets 39K extra images → pretraining data vs language effect not separated (Table III is the cleanest evidence).
  • Zero-shot (K = 0) protocol with a trained head is unclear; ablation rows (58–83%) measured in a different setting than the 95.06% headline.
  • Language = short adjective lists (language-only 60%); may discard geometry/slip cues; vision-based sensors and recognition only.

Two-minute summary

"Robots need touch; tactile cameras differ, so models fail across sensors. The authors use language as a sensor-agnostic teacher: a frozen BART embeds touch descriptions, a ViT is trained to match them with a temperature-softened KL loss (T 3.5) mixed 0.25/0.75 with cross-entropy, then frozen; only a tiny head is trained per task. On 39K relabelled DIGIT samples (32 classes) it reaches 95% at 100 shots, +13.3% average cross-sensor gain to GelSight, 98.8% on HCT, and beats vision-only using touch alone."

Viva cheat sheet · Paper 2page 5 of 6
P3

MARL sim2real transfer for multi-UAV delivery

Shi, Liu, Zhang, Zhou, Wang 2023, IEEE Trans. SMC: Systems 53(4) · Tongji / Monmouth

Workflow and modules

domain parameters μ from real task→AirSim + domain randomization→train R-MADDPG (20k ep)→test in sim (1000 ep)→direct transfer to F450/PX4 UAVs→fail? revise μ or policy
camera pixels→CNN detector (perception, swapped sim↔real)→o_pept boxes + o_loc (GPS)→MARL control (transferred)→(vx, vy, yaw) every 0.8 s

Definitions & equations

$$\text{NS Markov game }\langle N,S,A_i,T,O_i,r,\gamma,\mu\rangle,\quad T:S\times A_1\times\dots\times A_N\times\mu\to S,\quad R_i=\mathbb E\big[\textstyle\sum_t\gamma^tr^t\big]$$ $$\text{MADDPG: }\pi_i(o_i)=a_i,\ Q_i^\pi(x,c,z);\qquad \nabla_{\theta_i}J=\mathbb E[\nabla_{\theta_i}\pi_i(o_i)\nabla_{a_i}Q_i(x,a_1..a_N)|_{a_i=\pi_i(o_i)}]$$ $$\text{R-MADDPG: }\pi_i(o_i^t,y_i^t),\ y_i^t=y(h_i^t),\ h_i^t=(o_i^{t-1},o_i^{t-2},\dots);\quad Q_i(x^t,c^t,z^t,\mu,f^t)$$ $$r=\alpha r_{d1}+\beta r_{d2}+\omega r_{com};\ r_{d1}=\textstyle\sum_i(d_{1,i}^{t-1}-d_{1,i}^t);\ r_{d2}=-|d_2^{t-1}-d_2^t|\ (\text{or pun});\ r_{com}=\Gamma-t$$

Numbers to quote

HardwareF450 quadrotor, PX4, differential GPS, front+down mono cameras, Jetson Nano (4× A57, 128-core Maxwell)Obs / acto_loc(lx,ly) + o_pept(px1,py1,px2,py2); (vx, vy, yaw) continuousμ randomisedσl, σp, FP prob., flight height, D1, D2 (uniform per episode); GPS err 1–3 mNot randomisedmass, controller gains, wind (PX4 absorbs)Training20,000 ep · ≤30 steps · 0.8 s · lr 5e-3 · γ 0.99 · batch 128 ep · update /16 ep · Adam · ≈40 h/policyNetworkFC branch + (FC→LSTM) branch, concat, 2 FC, ReLU, 64 units, tanh actorEval40 ep every 270 training ep; 1000 test ep; real flightsPoliciesRMRand · MRand · RARand · RCRand · MADDPG (no DR)

Theory in one breath

Two nonstationarities: inherent (other agents learn; solved by CTDE critic) vs system (world changes via μ; solved by recurrent policy + DR). Proposition 1: under A1 (others = environment) and A2 (observation bijective), an RNN policy over history can infer $P_\mu$ hence μ (Glivenko–Cantelli), so recurrent CTDE MARL solves the game. Gaps: i.i.d. assumption, 30 steps, no learning guarantee.

Results in one breath

RMRand > MRand (memory helps under DR). RARand ≫ RCRand, MRand (memory in the actor is what matters; ≈ RMRand). MRand > MADDPG (DR helps even without memory). Sim: similar completion, RMRand ≈ RARand fastest. Real world: RMRand best completion and time; MADDPG without DR fails completely. Noise intuition: σ/√n (3 m → 1.34 m over 5 steps).

Critique (prepare calmly)

  • Two UAVs, no obstacles, fixed altitude, 2-D control: minimal "collective intelligence".
  • No seeds/variance/number of real flights; lr 5e-3 high for Adam; table numbers only in figures.
  • Proposition 1 informal; perception sim2real sidestepped by swapping detectors; dynamics (mass, wind) trusted to PX4.
  • Metaverse framing is conceptual; the substance is a MARL sim2real experiment (still rare on real multi-robot hardware).

Two-minute summary

"Two drones carry goods on ropes to two markers. Training in reality is unsafe, so they train in AirSim with domain randomization over six parameters so the real world is one sample of the simulated distribution. Policies are R-MADDPG: LSTM memory in actor and critic so hidden conditions can be inferred from history; perception is a separate CNN detector. The memory-based randomised policy transfers directly to real F450/PX4 drones with the best completion rate; the actor's memory matters more than the critic's; without randomization the policy fails."

Viva cheat sheet · Paper 3page 6 of 6