Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 8 min read

Section 2: The general principle of generative models

I am working through Mathematical Foundations of Generative AI, Prof. Prathosh AP's public lectures. The playlist is the spine. These notes are my deep dive on each section: a visual when the picture is the point, a deto

I am working through Mathematical Foundations of Generative AI, Prof. Prathosh AP's public lectures. The playlist is the spine. These notes are my deep dive on each section: a visual when the picture is the point, a detour when a prerequisite is doing real work, and the formula in my own words.

Previously: Section 1: The dataset and the problem.

Section 2: The general principle of generative models

Section 1 left me with a vague goal: estimate $P_x$ and learn to sample from it. This is the recipe that turns that goal into three moves. Every later model in the lectures is this recipe with different choices plugged in.

The statement I am carrying:

(i) Assume a parametric family on $P_x$, denoted $P_\theta$, represented using deep neural networks.
(ii) Define and estimate a divergence between $P_x$ and $P_\theta$.
(iii) Solve an optimisation problem over the parameters of $P_\theta$ to minimise that divergence.

The example that makes it concrete: draw $z \sim \mathcal{N}(0, I)$, pass it through a network $g_\theta : \mathcal{Z} \to \mathcal{X}$, and call the distribution of $x = g_\theta(z)$ by the name $P_\theta$. Then solve

$$
\theta^* = \arg\min_\theta \; \mathcal{D}(P_x \,|\, P_\theta)
$$

After that, sampling $z \sim \mathcal{N}(0, I)$ and computing $g_{\theta^}(z)$ gives samples from $P_{\theta^}$, which is as close to $P_x$ as this family could get.

Four prerequisites sit under that formula: the standard Gaussian $\mathcal{N}(0, I)$, what happens when a random variable is pushed through a function, what a divergence is, and the $\arg\min$ notation. Each detour is only as long as the step that needs it.

Intuition

I cannot see $P_x$. I build a machine with knobs, $\theta$, that turns plain random noise into data-shaped outputs. I keep turning the knobs until the machine's output distribution is as close as I can get to the real one. The closeness is a divergence. The turning is optimisation.

The load test, continued

The traffic generator from Section 1 takes random seeds and outputs synthetic requests. It has config knobs: endpoint weights, payload-size parameters, and so on.

Load-testing analogy Notation
Random seed fed into the generator $z \sim \mathcal{N}(0, I)$
The generator code $g_\theta$, a neural network
Its config knobs $\theta$, the network weights
The distribution of traffic it produces $P_\theta$
A score of how different synthetic traffic is from real logs the divergence $\mathcal{D}(P_x \,|\, P_\theta)$
Tuning knobs until that score is near zero $\theta^* = \arg\min_\theta \mathcal{D}$

The hard step is the score. I cannot see the true user behaviour. I only have the log. That is the difficulty the rest of the lectures keep returning to.

Visual

One dimension is enough to see the recipe. The noise is $z \sim \mathcal{N}(0, 1)$. The network is the simplest one I can write, $g_\theta(z) = a z + b$, so $\theta = (a, b)$. The green curve is the target. The blue curve is what this generator can produce.

Three things I want in front of me while I move the sliders:

  1. On the single bump, $a = 0.6$ and $b = 2$ drive the divergence to 0. The target is a bell centred at 2 with spread 0.6, and $x = 0.6z + 2$ is exactly that bell. The machine reproduces $P_x$.
  2. On the two bumps, the divergence never reaches 0, however I set $(a, b)$. The closest blue curve is a wide bell laid over both humps. A linear $g_\theta$ can only shift and stretch one bell. It cannot make two. That is why the principle asks for a deep network: $g_\theta$ has to be flexible enough that $P_\theta$ can get near $P_x$.
  3. The chart can draw the green curve, and it can compute the divergence, because I built the target in. In the real problem I do not get that picture. I only get samples. That is the first of the four open questions below.

Detours

The standard Gaussian $\mathcal{N}(0, I)$. $\mathcal{N}(\mu, \sigma^2)$ is the bell centred at $\mu$ with spread $\sigma$. Standard means $\mu = 0$ and $\sigma = 1$. In $k$ dimensions, $\mathcal{N}(0, I)$ is $k$ independent standard bells, one per coordinate. $I$ is the $k \times k$ identity covariance: each coordinate has variance 1, and no two coordinates are correlated.

It is the input noise because every library can sample it, and it has no structure of its own. Any structure in the output has to come from $g_\theta$. So $z \sim \mathcal{N}(0, I)$ is a random seed: easy to draw, carrying nothing about images.

Pushing a random variable through a function. If $z$ is random and $x = g(z)$, then $x$ is random too, with a new distribution determined by $g$.

The simplest case: if $z \sim \mathcal{N}(0, 1)$, then $x = az + b \sim \mathcal{N}(b, a^2)$. Multiplying by $a$ stretches the spread. Adding $b$ slides the centre. That is what the sliders above are doing.

A deep, nonlinear $g$ can bend the output into more than one hump, or into a thin set such as the distribution of faces. The output distribution is different from the distribution of $z$, and it depends on $g_\theta$. I name that output distribution $P_\theta$. For a deep network I usually cannot write a formula for it. I only know how to sample it: draw $z$, run $g_\theta$.

Divergence, which is weaker than a distance. $\mathcal{D}(P \,|\, Q)$ scores how different two distributions are. The two properties I need:

$$
\mathcal{D}(P_x \,|\, P_\theta) \ge 0, \qquad \mathcal{D}(P_x \,|\, P_\theta) = 0 \iff P_x = P_\theta
$$

It need not be symmetric. $\mathcal{D}(P \,|\, Q)$ may differ from $\mathcal{D}(Q \,|\, P)$. That is why the notation uses $|$ rather than a comma. The KL divergence in the chart is one example. The next part of the lectures builds a whole family of them, the f-divergences.

Those two properties are what make "minimise the divergence" a sensible goal. The smallest possible value is 0, and it is reached only when the model matches the data.

$\arg\min$. $\min_\theta f(\theta)$ is the smallest value of $f$. $\arg\min_\theta f(\theta)$ is the input $\theta$ that achieves it. For $f(\theta) = (\theta - 3)^2 + 1$, the min is 1 and the argmin is 3. I want the knob settings, not the score, so the principle uses argmin.

The formula

$$
\theta^* = \arg\min_\theta \; \mathcal{D}\big(P_x \,|\, P_\theta\big), \qquad P_\theta := \text{distribution of } g_\theta(z),\; z \sim \mathcal{N}(0, I)
$$

Read aloud: $\theta^*$ is the setting of the network weights that makes the divergence between the true data distribution and the distribution of the network's outputs as small as possible.

Symbol What it is Type / shape Role
$z$ input noise vector in $\mathbb{R}^k$, usually $k \ll d$ the random seed
$\mathcal{N}(0, I)$ standard Gaussian distribution on $\mathbb{R}^k$ easy-to-sample source of randomness
$g_\theta$ generator function $\mathbb{R}^k \to \mathbb{R}^d$, a neural net reshapes noise into data
$\theta$ generator parameters a vector of weights the knobs
$P_\theta$ model distribution distribution on $\mathbb{R}^d$, implicit what the generator produces
$P_x$ true distribution distribution on $\mathbb{R}^d$, unknown the target
$\mathcal{D}(\cdot \,|\, \cdot)$ divergence two distributions to a scalar $\ge 0$ the "how far apart" score
$\theta^*$ optimal parameters same shape as $\theta$ the trained generator

Once training is done, sampling is: draw a fresh $z \sim \mathcal{N}(0, I)$, compute $g_{\theta^}(z)$. That is a sample from $P_{\theta^} \approx P_x$. I have learned to sample from $P_x$ without writing it down.

A worked example small enough to do by hand

Take $k = d = 1$ and $g_\theta(z) = 2z + 3$, so $\theta = (a, b) = (2, 3)$.

  • Draw three noise values: $z = -1, 0, 1$.
  • Outputs: $g_\theta(-1) = 1$, $g_\theta(0) = 3$, $g_\theta(1) = 5$.
  • By the pushforward above, all outputs come from $P_\theta = \mathcal{N}(3, 2^2)$. Centre 3, spread 2, which matches these samples.

For the divergence, a two-outcome world is enough to compute by hand. Let $P_x$ be a fair coin, $(0.5, 0.5)$, and use KL, $\sum_i P_x(i) \ln \frac{P_x(i)}{P_\theta(i)}$. KL is formally the next part of the lectures. I only need it here to see numbers move.

  • If the model says $P_\theta = (0.9, 0.1)$: $0.5 \ln\frac{0.5}{0.9} + 0.5 \ln\frac{0.5}{0.1} = 0.5(-0.588) + 0.5(1.609) \approx 0.511$.
  • If the model says $P_\theta = (0.6, 0.4)$: $0.5 \ln\frac{0.5}{0.6} + 0.5 \ln\frac{0.5}{0.4} = 0.5(-0.182) + 0.5(0.223) \approx 0.020$.
  • If the model says $P_\theta = (0.5, 0.5)$: both terms are $0.5 \ln 1 = 0$, so the divergence is exactly 0.

As the model gets closer to the truth, the score falls toward 0, and it hits 0 only at a perfect match. Optimisation is moving $\theta$ in the direction that lowers this number.

Four questions this recipe does not answer

The rest of the lectures is the attempt to answer these.

Question Why it is hard Where I expect the answer
How do I compute the divergence without formulas for $P_x$ or $P_\theta$? I only have samples of both Write the divergence as expectations, then estimate them with sample averages
Which divergence? Different choices behave differently. The wide bell over two bumps is one failure mode f-divergences first, Wasserstein later
How do I choose $g_\theta$, and so $P_\theta$? Flexible enough, and still trainable GANs, VAEs, diffusion, transformers
How do I solve the optimisation? Millions of parameters, often a min-max game Gradient descent, then its adversarial variants

Where the later models sit

Model Choice of $P_\theta$ Divergence Optimisation
GAN $g_\theta(z)$, implicit Jensen–Shannon, an f-divergence min-max with a discriminator
WGAN $g_\theta(z)$, implicit Wasserstein min-max with a 1-Lipschitz critic
VAE latent model $\int P_\theta(x \mid z)\, p(z)\, dz$ KL, through the ELBO maximise the ELBO
DDPM reverse chain of Gaussians KL, through the ELBO regression on the added noise
Autoregressive / transformer $\prod_t P_\theta(x_t \mid x_{<t})$ KL, which is maximum likelihood cross-entropy

One link I want early: minimising $\mathrm{KL}(P_x \,|\, P_\theta)$ is the same as maximum likelihood, because $\arg\min \mathrm{KL}(P_x \,|\, P_\theta) = \arg\max \mathbb{E}[\log P_\theta(x)]$. That is why VAEs, diffusion models, and language models can all train on log-likelihood losses and still be this same recipe.

Questions I want to be able to answer:

  1. Why can no setting of $(a, b)$ match the two-bump target, and what has to change about $g_\theta$?
  2. After training, how do I produce 10 new samples, in one line?
  3. If $\mathcal{D}(P_x \,|\, P_\theta) = 0$, what follows? What if the divergence is small but not zero?
πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.