<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://abdgafartunde.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://abdgafartunde.github.io/" rel="alternate" type="text/html" /><updated>2026-10-05T14:22:04+08:00</updated><id>https://abdgafartunde.github.io/feed.xml</id><title type="html">A.T. Tiamiyu</title><subtitle>Postdoctoral Researcher in Computational and Applied Mathematics</subtitle><author><name>Abd&apos;gafar Tunde Tiamiyu</name><email>abdgafartunde@yahoo.com</email></author><entry><title type="html">Primal-Dual Methods for Convex Optimization</title><link href="https://abdgafartunde.github.io/blog/2026/10/05/primal-dual-methods-convex-optimization/" rel="alternate" type="text/html" title="Primal-Dual Methods for Convex Optimization" /><published>2026-10-05T00:00:00+08:00</published><updated>2026-10-05T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/10/05/primal-dual-methods-convex-optimization</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/10/05/primal-dual-methods-convex-optimization/"><![CDATA[<p>Gradient descent is the workhorse of smooth optimization, and much of machine learning runs on it or variants of it. But the objectives that arise naturally in imaging and inverse problems are often not smooth. Total variation regularization, $\ell^1$ norms, indicator functions of convex sets: these are non-differentiable, and gradient descent simply does not apply.</p>

<p>The standard workaround is to replace the non-smooth term with a smooth approximation. This works, but it introduces a smoothing parameter that must be tuned and that alters the problem being solved. For applications where the non-smooth structure matters (edge-preserving regularization, exact sparsity, or hard constraints), this is not ideal.</p>

<p>Primal-dual methods offer an alternative: algorithms designed specifically for non-smooth convex objectives that converge to the exact solution of the original problem. They are the basis of the optimization algorithms I use in my own EIT reconstruction work, and understanding their structure is useful for anyone working on imaging or variational inverse problems.</p>

<h2 id="the-setting-composite-convex-minimization">The Setting: Composite Convex Minimization</h2>

<p>The problems we want to solve have the form</p>

\[\min_{u \in \mathcal{X}} \; F(Ku) + G(u),\]

<p>where $\mathcal{X}$ and $\mathcal{Y}$ are Hilbert spaces, $K: \mathcal{X} \to \mathcal{Y}$ is a bounded linear operator, and $F: \mathcal{Y} \to \mathbb{R} \cup {+\infty}$ and $G: \mathcal{X} \to \mathbb{R} \cup {+\infty}$ are proper, convex, lower semicontinuous functions.</p>

<p>For variational reconstruction with total variation, $u$ is the unknown image, $G(u) = \frac{1}{2}|Au - b|^2$ is the data-fidelity term, $K = \nabla$ is the gradient operator, and $F(z)=\alpha|z|_1$. Thus $F(Ku)=\alpha|\nabla u|_1$. The operator composition and nonsmooth regularizer make ordinary gradient descent inapplicable.</p>

<h2 id="the-saddle-point-reformulation">The Saddle-Point Reformulation</h2>

<p>The key idea of primal-dual methods is to reformulate the composite minimization as a <strong>saddle-point problem</strong>. Using the Fenchel conjugate $F^<em>$ (defined as $F^</em>(y) = \sup_z \langle z, y \rangle - F(z)$), we have</p>

\[F(Ku) = \sup_{p \in \mathcal{Y}} \; \langle Ku, p \rangle - F^*(p).\]

<p>Substituting:</p>

\[\min_{u} \; F(Ku) + G(u) = \min_u \max_p \; \langle Ku, p \rangle - F^*(p) + G(u).\]

<p>This saddle-point formulation introduces a <strong>dual variable</strong> $p$. In a Hilbert space, the Riesz representation identifies $\mathcal{Y}$ with its dual, which is why the same space appears in the supremum.</p>

<p>For $F(z)=\alpha|z|<em>1$, the Fenchel conjugate is the indicator function $F^*(p)=\iota</em>{|p|_\infty\leq\alpha}(p)$. The dual variable is therefore constrained to the corresponding pointwise dual-norm ball. For isotropic TV, the norm at each pixel is the Euclidean norm of the discrete gradient vector.</p>

<h2 id="the-primal-dual-hybrid-gradient-algorithm">The Primal-Dual Hybrid Gradient Algorithm</h2>

<p>The <strong>primal-dual hybrid gradient</strong> (PDHG) algorithm, introduced by Chambolle and Pock (2011), solves the saddle-point problem by alternating proximal steps in the primal and dual variables:</p>

\[\begin{aligned}
p^{k+1} &amp;= \mathrm{prox}_{\sigma F^*}\!\left(p^k + \sigma K \bar{u}^k\right) \\
u^{k+1} &amp;= \mathrm{prox}_{\tau G}\!\left(u^k - \tau K^* p^{k+1}\right) \\
\bar{u}^{k+1} &amp;= u^{k+1} + \theta(u^{k+1} - u^k)
\end{aligned}\]

<p>where $\sigma, \tau &gt; 0$ are step sizes, $\theta \in [0, 1]$ is an over-relaxation parameter (typically $\theta = 1$), and $\bar{u}$ is an extrapolated (“looking ahead”) version of $u$.</p>

<p>The proximal operator $\mathrm{prox}_{\lambda f}(v) = \arg\min_u \frac{1}{2}|u - v|^2 + \lambda f(u)$ is the key computational primitive. For many common functions, it has a closed-form expression:</p>

<ul>
  <li><em>*$F^</em> = \iota_{|p|<em>\infty \leq \alpha}$** (total variation): for isotropic TV, $\mathrm{prox}</em>{\sigma F^*}(q) = q / \max(1, \lvert q \rvert/\alpha)$ pointwise, where $\lvert q \rvert$ is the Euclidean norm of the local gradient vector. The result is the projection onto the dual ball.</li>
  <li><strong>$G(u) = \frac{1}{2}|Au - b|^2$</strong> (least-squares fidelity): the proximal operator is $(I + \tau A^\top A)^{-1}(u + \tau A^\top b)$, which requires solving a linear system. For small problems this is direct; for large problems this inner solve is done iteratively (often with CG, linking back to the previous post).</li>
  <li><strong>$G = \iota_C$</strong> (indicator of a convex set $C$): $\mathrm{prox}$ is projection onto $C$.</li>
</ul>

<h2 id="convergence">Convergence</h2>

<p>Assume that the saddle-point problem has a solution. For the standard Chambolle-Pock update with $\theta=1$, the condition $\sigma\tau|K|^2&lt;1$ gives convergence, and the ergodic primal-dual gap decays as $O(1/N)$ under the usual finite-dimensional assumptions. When the primal or dual objective is uniformly convex, modified accelerated step-size rules can improve the corresponding rate to $O(1/N^2)$. Linear rates require stronger assumptions on both sides.</p>

<p>The algorithm requires only:</p>
<ul>
  <li>Operator products with $K$ and $K^*$, without forming a dense matrix</li>
  <li>Proximal operators of $F^*$ and $G$ separately</li>
  <li>No joint evaluation of $F$ and $G$ together</li>
</ul>

<p>This separation is the key to handling composite non-smooth objectives: by working with $F$ and $G$ separately, PDHG avoids having to differentiate through their composition.</p>

<h2 id="extensions-and-variants">Extensions and Variants</h2>

<p><strong>Multiple nonsmooth terms.</strong> An objective containing data fidelity, TV, and a positivity constraint can be lifted to a product space with one operator block per composite term. Primal-dual splitting, Douglas-Rachford splitting, and ADMM provide different ways to handle the resulting structure; their updates and convergence assumptions are not interchangeable.</p>

<p><strong>Strongly convex acceleration.</strong> When the primal or dual objective is uniformly convex, PDHG can use iteration-dependent primal and dual step sizes to obtain an accelerated $O(1/N^2)$ rate for the relevant error measure. Without strong convexity, the general ergodic gap rate remains $O(1/N)$.</p>

<p><strong>Nonlinear extensions.</strong> For non-convex problems (which arise in nonlinear inverse problems), primal-dual methods can still be applied as heuristics or extended via linearization strategies. In my EIT reconstruction work, I use a linearize-and-solve approach: the nonlinear forward operator is linearized around the current iterate, and the resulting quadratic subproblem is solved with PDHG. This is a Gauss-Newton outer loop with PDHG as the inner solver.</p>

<h2 id="why-this-matters-for-inverse-problems">Why This Matters for Inverse Problems</h2>

<p>In regularized inversion, the choice of regularizer is a modelling choice: it encodes prior belief about the structure of the unknown. Total variation says the unknown is piecewise smooth with sharp edges. The $\ell^1$ norm of coefficients in a dictionary says the unknown is sparse in that dictionary. These structural priors frequently reflect genuine knowledge about the problem, and they happen to be non-smooth.</p>

<p>Primal-dual methods handle nonsmooth terms without replacing them by smooth approximations. Under their convergence assumptions, the iterates approach a minimizer of the regularized objective. Implementations use operator evaluations, adjoints, and separate proximal maps or proximal subproblem solves.</p>

<p>The combination of variational regularization with primal-dual algorithms is, in my view, one of the most useful and underappreciated toolsets in computational mathematics. It is worth understanding well.</p>

<p>The convergence statements above follow Chambolle and Pock, <a href="https://doi.org/10.1007/s10851-010-0251-1">“A first-order primal-dual algorithm for convex problems with applications to imaging”</a>.</p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Mathematics" /><category term="Optimization" /><category term="Inverse Problems" /><category term="Scientific Computing" /><summary type="html"><![CDATA[How saddle-point formulations and primal-dual splitting algorithms handle the non-smooth objectives that arise in imaging and inverse problems; total variation regularization requires going beyond gradient descent.]]></summary></entry><entry><title type="html">Electrical Impedance Tomography in Clinical Practice: From Research to Bedside</title><link href="https://abdgafartunde.github.io/blog/2026/09/21/eit-clinical-applications/" rel="alternate" type="text/html" title="Electrical Impedance Tomography in Clinical Practice: From Research to Bedside" /><published>2026-09-21T00:00:00+08:00</published><updated>2026-09-21T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/09/21/eit-clinical-applications</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/09/21/eit-clinical-applications/"><![CDATA[<p>For most of my PhD, electrical impedance tomography was a mathematical object: a nonlinear inverse problem with partial boundary data, an ill-posed mapping from boundary voltages to interior conductivity distributions, a setting in which to develop and test regularization algorithms. That framing is accurate, but it is incomplete.</p>

<p>EIT is also a clinical monitoring technology. Commercial devices are used in intensive care and studied as tools for assessing regional ventilation and supporting ventilator adjustment. Evidence for physiological monitoring is substantial, but evidence that EIT-guided care improves patient-centred outcomes remains limited. Understanding this distinction is directly relevant to the mathematical and computational problems.</p>

<h2 id="what-eit-does">What EIT Does</h2>

<p>EIT estimates changes in internal electrical conductivity from voltage measurements made at the surface. A belt of electrodes is attached around the patient’s thorax. Small alternating currents are applied through selected electrodes while voltages are recorded at the others. Cycling through many configurations produces a vector of boundary measurements from which a regularized conductivity-change image is computed.</p>

<p>The conductivity of biological tissue varies by type: air-filled lung tissue is highly resistive, while blood and soft tissues are more conductive. As the lungs inflate and deflate, the regional distribution of air changes, and so does the regional impedance. Clinical systems can produce tens of images per second, providing continuous bedside monitoring of ventilation distribution.</p>

<h2 id="the-clinical-problem-it-solves">The Clinical Problem It Solves</h2>

<p>Patients in intensive care units who cannot breathe independently are supported by mechanical ventilators that deliver air under positive pressure. Getting the ventilation parameters right (the tidal volume, the respiratory rate, the positive end-expiratory pressure) is critical and difficult.</p>

<p>Set the pressure too low, and under-aerated lung regions collapse (atelectasis) and are re-recruited with each breath, causing repetitive trauma. Set it too high, and over-distended lung regions are damaged (volutrauma). In patients with acute respiratory distress syndrome (ARDS), where lung compliance is severely impaired and uneven, the optimal ventilation settings vary substantially between patients and can change over time as the condition evolves.</p>

<p>The problem is that clinicians have had few tools for monitoring regional lung function at the bedside. Conventional chest X-ray shows global lung structure but not function; CT is high-resolution but requires transporting a critically ill patient to a scanner and exposes them to radiation; pulse oximetry measures oxygen saturation but cannot localize which regions of the lung are contributing.</p>

<p>EIT addresses part of this monitoring gap. It provides continuous, radiation-free bedside information about regional ventilation and changes during interventions such as PEEP trials or prone positioning. Clinicians can use this information alongside pressure, flow, gas-exchange, and imaging data. EIT does not by itself determine an optimal ventilator setting, and high-quality evidence for improved clinical outcomes is still developing.</p>

<h2 id="clinical-deployment-and-evidence">Clinical Deployment and Evidence</h2>

<p>The Dräger PulmoVista 500 is one documented commercial example. It uses a thoracic belt with 16 integrated electrodes and displays regional ventilation information continuously at the bedside. Other research and commercial systems use related measurement principles, although reconstruction methods, interfaces, and regulatory approvals differ.</p>

<p>International consensus work now covers EIT acquisition, processing, terminology, and use in adult critical care. The evidence is strongest for monitoring regional ventilation and physiological responses. A 2025 evidence-based consensus found that only a small fraction of its clinical recommendations were supported by high-level evidence and noted the absence of multicentre randomized trials for many proposed uses. Clinical deployment should therefore be described as established monitoring with an evolving outcomes evidence base, not as a universally validated replacement for conventional imaging or ventilator protocols.</p>

<h2 id="what-the-mathematics-currently-provides">What the Mathematics Currently Provides</h2>

<p>Clinical thoracic EIT commonly uses fast linearized reconstruction of the difference between current and reference measurements. This produces images of conductivity <em>change</em>, not absolute conductivity. Spatial resolution is low compared with CT or MRI, but temporal resolution is high.</p>

<p>This is a deliberate engineering choice. For the clinical application of tracking lung ventilation changes, high temporal resolution matters more than spatial resolution; the clinician wants to know whether the left lung or right lung is better ventilated, not a precise anatomical map. The simple reconstruction algorithm is fast enough for real-time display and robust enough for clinical use.</p>

<p>The research frontier is different. Current academic work addresses:</p>

<ul>
  <li><strong>Absolute EIT</strong> (reconstructing conductivity, not just changes): this requires accurate electrode models, precise contact impedance characterization, and solving a much harder inverse problem, but it would enable tissue characterization rather than just functional monitoring.</li>
  <li><strong>3D reconstruction</strong>: commercial devices image a 2D cross-section; the lung is three-dimensional, and out-of-plane currents limit accuracy. 3D EIT requires more electrodes, more data, and far more expensive reconstruction.</li>
  <li><strong>Partial boundary data</strong>: in some clinical settings (post-mastectomy, obese patients, surgical wounds), the full electrode ring cannot be placed. Reconstruction from partial boundary data is a mathematically harder problem with important clinical applications.</li>
  <li><strong>Learned regularizers</strong>: data-driven approaches that use anatomical priors from CT databases to improve spatial resolution and robustness.</li>
</ul>

<p>The last three years of my own PhD research sat at the intersection of partial boundary data, regularization theory, and learned regularizers, motivated precisely by the gap between what current clinical devices do and what is mathematically possible.</p>

<h2 id="the-gap-between-current-and-potential">The Gap Between Current and Potential</h2>

<p>EIT avoids ionizing radiation, patient transport, large magnets, and cryogenic infrastructure. It can operate continuously at the bedside. These properties make it attractive for physiological monitoring where CT or MRI is unavailable or impractical, but EIT does not provide their anatomical resolution or diagnostic scope.</p>

<p>The limitation is image quality. Resolution is low, and absolute conductivity imaging remains clinically unreliable. Closing this gap is an algorithmic and mathematical problem as much as a hardware problem. Better reconstruction algorithms (ones that exploit prior information more effectively, handle partial data more gracefully, and are robust to the inevitable imperfections in clinical electrode placement) would translate directly into clinical benefit.</p>

<p>That translation is not guaranteed, but it is tractable. The mathematics is hard, the clinical need is genuine, and the computational tools available now are substantially more powerful than they were when the field started. That seems like a reasonable place to be working.</p>

<p><em>For implementations of EIT reconstruction algorithms, see the <a href="https://github.com/abdgafartunde/eit-reconstruction-toolkit">eit-reconstruction-toolkit GitHub repository</a> (to be updated).</em></p>

<p>Clinical context and evidence: the <a href="https://doi.org/10.1186/s13054-024-05173-x">2024 international expert report on adult ICU EIT</a>, the <a href="https://doi.org/10.1016/j.eclinm.2025.103575">2025 evidence-based consensus</a>, and the <a href="https://www.draeger.com/en-us_ca/Products/PulmoVista-500">PulmoVista 500 product documentation</a>.</p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Inverse Problems" /><category term="Medical Imaging" /><category term="EIT" /><category term="Applied Mathematics" /><summary type="html"><![CDATA[How EIT supports continuous bedside monitoring of regional lung ventilation, what current evidence supports, and which reconstruction problems remain open.]]></summary></entry><entry><title type="html">Krylov Methods: Solving Large Linear Systems Without Forming the Matrix</title><link href="https://abdgafartunde.github.io/blog/2026/09/07/krylov-methods-iterative-solvers/" rel="alternate" type="text/html" title="Krylov Methods: Solving Large Linear Systems Without Forming the Matrix" /><published>2026-09-07T00:00:00+08:00</published><updated>2026-09-07T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/09/07/krylov-methods-iterative-solvers</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/09/07/krylov-methods-iterative-solvers/"><![CDATA[<p>Almost every large-scale computation in scientific computing eventually reduces to solving a linear system $Ax = b$. Whether the system comes from discretizing a PDE, from the normal equations of a least-squares problem, from the linearization step in a Newton iteration, or from the regularized inversion of a forward operator: the core computational task is the same.</p>

<p>For dense systems, Gaussian elimination requires $O(n^3)$ operations and $O(n^2)$ storage. Sparse direct methods can be much cheaper and remain effective for many moderate-sized problems, but fill-in can make their memory and computational costs prohibitive for large three-dimensional PDE systems. Problems in image reconstruction or seismic inversion may have millions of unknowns, so matrix-free iterative methods are often the practical choice.</p>

<p>Among iterative solvers, <strong>Krylov subspace methods</strong> are widely used because they require matrix-vector products and can exploit sparse or matrix-free operators. Their performance still depends on the matrix structure, stopping criterion, and preconditioner.</p>

<h2 id="the-krylov-subspace">The Krylov Subspace</h2>

<p>Given an initial guess $x_0$ with residual $r_0=b-Ax_0$, the <strong>Krylov subspace</strong> of order $k$ is</p>

\[\mathcal{K}_k(A, r_0) = \operatorname{span}\{r_0, \, Ar_0, \, A^2 r_0, \, \ldots, \, A^{k-1} r_0\}.\]

<p>A Krylov method seeks an approximate solution in the affine space $x_0+\mathcal{K}_k(A,r_0)$. The meaning of “best” depends on the method: CG minimizes the energy norm of the error, while GMRES minimizes the Euclidean norm of the residual.</p>

<p>The key insight is that each successive Krylov subspace can be built from the previous one by multiplying by $A$, one matrix-vector product per iteration. You never need to form or store $A$ explicitly; you only need to compute matrix-vector products $A v$ for given vectors $v$. For sparse matrices arising from PDE discretizations, each product takes $O(n)$ or $O(n \log n)$ operations. This is what makes Krylov methods scalable.</p>

<h2 id="the-conjugate-gradient-method">The Conjugate Gradient Method</h2>

<p>For symmetric positive definite (SPD) matrices, the <strong>conjugate gradient</strong> (CG) method is optimal among Krylov methods. It minimizes the $A$-norm of the error over the Krylov subspace:</p>

\[x_k = \arg\min_{x \in x_0+\mathcal{K}_k(A,r_0)} \|x - x^*\|_A, \quad \|v\|_A^2 = v^\top A v.\]

<p>The iteration is:</p>

\[\begin{aligned}
r_0 &amp;= b - Ax_0, \quad p_0 = r_0 \\
\alpha_k &amp;= \frac{r_k^\top r_k}{p_k^\top A p_k} \\
x_{k+1} &amp;= x_k + \alpha_k p_k \\
r_{k+1} &amp;= r_k - \alpha_k A p_k \\
\beta_k &amp;= \frac{r_{k+1}^\top r_{k+1}}{r_k^\top r_k} \\
p_{k+1} &amp;= r_{k+1} + \beta_k p_k
\end{aligned}\]

<p>Each iteration requires one matrix-vector product $Ap_k$, two inner products, and a few vector additions. For a system of size $n$, the exact solution is reached in at most $n$ iterations in exact arithmetic. The standard worst-case bound implies $O(\sqrt{\kappa}\log(1/\varepsilon))$ iterations to reduce the energy-norm error by a factor $\varepsilon$, where $\kappa = \lambda_{\max}/\lambda_{\min}$. For preconditioned CG, $\kappa$ refers to the appropriately preconditioned SPD operator.</p>

<h2 id="convergence-and-the-eigenvalue-distribution">Convergence and the Eigenvalue Distribution</h2>

<p>The convergence of CG depends on the spectrum of $A$. If $A$ has $m$ distinct eigenvalues, CG converges in at most $m$ steps in exact arithmetic. More practically, if the eigenvalues are clustered into a few groups, CG can converge rapidly because a polynomial of low degree can be small on the eigenvalue clusters simultaneously.</p>

<p>This geometric picture (that Krylov methods implicitly find a polynomial that is small on the spectrum of $A$) unifies all Krylov methods. The differences between methods (CG, MINRES, GMRES, BiCGSTAB) correspond to different choices of what “small” means and what structure the polynomial is required to respect.</p>

<h2 id="non-symmetric-systems-gmres">Non-Symmetric Systems: GMRES</h2>

<p>When $A$ is not symmetric, CG is not applicable. The most widely used alternative is <strong>GMRES</strong> (generalized minimal residual method), which minimizes the 2-norm of the residual over the Krylov subspace:</p>

\[x_k = \arg\min_{x \in x_0+\mathcal{K}_k(A,r_0)} \|b - Ax\|_2.\]

<p>GMRES builds an orthonormal basis for the Krylov subspace using the Arnoldi process (a Gram-Schmidt orthogonalization against all previous basis vectors), then solves a small least-squares problem in this basis. The price is memory: GMRES stores all $k$ basis vectors, so restarting (GMRES(m)) is often used in practice to bound memory use.</p>

<p>GMRES is commonly used for nonsymmetric systems arising from convection-diffusion equations, linearized fluid equations, and nonsymmetric Jacobian systems.</p>

<h2 id="preconditioning">Preconditioning</h2>

<p>For SPD systems, condition number and eigenvalue clustering provide useful CG convergence information. For nonsymmetric and nonnormal systems, the condition number alone is insufficient to predict GMRES convergence; spectral geometry and nonnormality also matter. Many PDE-derived systems become harder to solve as the mesh is refined, making unpreconditioned iterations impractically slow.</p>

<p><strong>Preconditioning</strong> transforms the system into one that is easier for the selected iterative method. Given a nonsingular preconditioner $P \approx A$, the left-preconditioned system $P^{-1}Ax = P^{-1}b$ has the same solution. A useful preconditioner is cheap to apply and improves the spectral or field-of-values properties relevant to convergence; a smaller condition number is not guaranteed for every nonsymmetric problem.</p>

<p>Common preconditioners include:</p>

<ul>
  <li><strong>Incomplete LU (ILU):</strong> a sparse approximate factorization of $A$, dropping fill-in below a threshold. Cheap to apply, often effective for mildly ill-conditioned problems.</li>
  <li><strong>Multigrid:</strong> a hierarchical solver that operates on a sequence of coarser and coarser grids. For elliptic PDEs, optimal multigrid preconditioners give $\kappa(P^{-1}A) = O(1)$ independently of the mesh size; this is essentially optimal and is the reason large-scale CFD and structural analysis codes use multigrid.</li>
  <li><strong>Block preconditioners:</strong> for saddle-point systems (e.g., the linearized Navier-Stokes equations, or regularized least-squares systems in inverse problems), block-diagonal or block-triangular preconditioners exploit the block structure.</li>
</ul>

<p>The design of good preconditioners for specific problem classes is itself a significant research area.</p>

<h2 id="krylov-methods-in-inverse-problems">Krylov Methods in Inverse Problems</h2>

<p>In computational inverse problems, Krylov methods appear in several places:</p>

<p><strong>Regularized normal equations.</strong> Tikhonov regularization of a linear inverse problem $Au \approx b$ leads to the normal equations $(A^\top A + \alpha L^\top L)u = A^\top b$. For $\alpha&gt;0$, this matrix is SPD when $\ker(A)\cap\ker(L)={0}$. In that case CG applies. The matrix need not be formed explicitly; products with $A$, $A^\top$, $L$, and $L^\top$ are sufficient.</p>

<p><strong>Iterative regularization.</strong> CG applied to a linear inverse problem (without explicit regularization) has a regularizing effect: early stopping acts as regularization, since the Krylov subspace captures the dominant singular components of $A$ before the noise-dominated ones. This is the basis of the CGLS and LSQR algorithms.</p>

<p><strong>Inexact Newton methods.</strong> In nonlinear inverse problems, each Newton step requires solving a linearized system. CG is appropriate for SPD Gauss-Newton or regularized normal equations, while GMRES is used for nonsymmetric linearizations. Truncating the inner solve according to a forcing rule can reduce computation while preserving outer convergence.</p>

<p><strong>Operator preconditioning.</strong> In function-space formulations of inverse problems (the Bayesian framework, or the Tikhonov formulation in Hilbert spaces), the preconditioner is typically the prior covariance or regularization operator. Applying it requires solving another PDE or working with the SPDE approximation to the GP prior, and this inner solve is also done iteratively.</p>

<p>The combination of Krylov methods with good preconditioners is what makes large-scale inverse problems computationally feasible. Without scalable iterative solvers, the algorithms I work on in EIT reconstruction and PDE-constrained optimization would be limited to toy problems.</p>

<p>For algorithm descriptions and guidance on matching methods to matrix structure, see Barrett et al., <a href="https://www.netlib.org/templates/templates.pdf"><em>Templates for the Solution of Linear Systems</em></a>.</p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Mathematics" /><category term="Scientific Computing" /><category term="Numerical Methods" /><summary type="html"><![CDATA[An introduction to Krylov subspace methods: the iterative solvers that make large-scale scientific computing possible, and why they work so well for the linear systems that arise in inverse problems and PDE discretizations.]]></summary></entry><entry><title type="html">Full Waveform Inversion: The Mathematics Behind Finding Oil</title><link href="https://abdgafartunde.github.io/blog/2026/08/24/full-waveform-inversion/" rel="alternate" type="text/html" title="Full Waveform Inversion: The Mathematics Behind Finding Oil" /><published>2026-08-24T00:00:00+08:00</published><updated>2026-08-24T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/08/24/full-waveform-inversion</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/08/24/full-waveform-inversion/"><![CDATA[<p>When a seismic survey vessel crosses the ocean, it tows an array of airguns that fire compressed air into the water at regular intervals. The acoustic waves travel down through the water, penetrate the seafloor, and propagate through kilometers of sediment and rock. Where the rock properties change (at the boundary between a shale layer and a sandstone reservoir, for instance, or beneath a dome of salt), some energy is reflected back. Hydrophones behind the vessel record these reflections as a function of time.</p>

<p>The question the geophysicist wants to answer is the inverse problem: given the recorded waveforms, what are the subsurface rock properties? More precisely, what is the velocity field $v(x)$ that governs how fast acoustic waves travel at each point in the subsurface? The velocity field encodes the lithology (the rock types and their spatial arrangement) and is the primary input for identifying where hydrocarbons might be trapped.</p>

<p>This inverse problem is called <strong>full waveform inversion</strong> (FWI). It is one of the most computationally demanding inverse problems in industrial use, and its mathematical structure is a direct application of the ideas in PDE-constrained optimization and the adjoint method.</p>

<h2 id="the-forward-problem">The Forward Problem</h2>

<p>The forward model is the acoustic wave equation. For a scalar pressure field $u(x, t)$ driven by a source $f(x, t)$:</p>

\[\frac{1}{v(x)^2} \frac{\partial^2 u}{\partial t^2} - \nabla^2 u = f(x, t),\]

<p>with appropriate initial and boundary conditions. Given the velocity model $v(x)$ and the source signature $f$, this PDE can be solved numerically (using finite differences or finite elements) to predict the wavefield $u$ and, by restriction to the receiver locations, the predicted seismograms $u_{\text{pred}}$.</p>

<p>The forward solve is expensive. For a 3D survey, the domain can be tens of kilometers in each direction, the wavefield must be evolved over several seconds of simulated time, and the spatial and temporal discretization must be fine enough to resolve the shortest wavelengths in the source signal. A single forward solve can take minutes on a large compute cluster, and a 3D FWI run requires thousands of forward solves.</p>

<h2 id="the-inverse-problem">The Inverse Problem</h2>

<p>FWI is formulated as a least-squares minimization:</p>

\[\min_{v} \; \mathcal{J}(v) = \frac{1}{2} \sum_{s} \| P_su_s(v) - d_s \|^2,\]

<p>where the sum is over seismic sources $s$, $u_s(v)$ is the simulated wavefield, $P_s$ restricts that wavefield to the receiver locations and sampled times, and $d_s$ is the observed seismogram.</p>

<p>This is a nonlinear, large-scale optimization problem. The state variable (the wavefield) satisfies a PDE that depends on the parameter $v$. Computing the gradient $\nabla_v \mathcal{J}$ by finite differences would require one forward solve per parameter, and $v$ is discretized on a grid with millions of nodes. This is computationally impossible.</p>

<p>The solution is the <strong>adjoint method</strong>. For each source $s$, the gradient can be computed with one forward solve to obtain $u_s(v)$ and one adjoint solve driven by the receiver residual $P_su_s(v)-d_s$, injected into the wave equation through $P_s^*$. The gradient contribution is a time integral of products of the forward and adjoint wavefields. The cost is therefore proportional to the number of sources, not the number of model parameters.</p>

<p>This is why adjoint-state differentiation is the standard scalable approach in FWI. Reverse-mode automatic differentiation computes the same mathematical adjoint, although its implementation and memory requirements may differ.</p>

<h2 id="why-it-is-hard">Why It Is Hard</h2>

<p>FWI is a highly nonlinear, non-convex optimization problem. The objective function has many local minima, and gradient-based methods will converge to a local minimum that may be geologically meaningless unless the starting model is sufficiently close to the true solution. The “cycle-skipping” problem (where the predicted and observed waveforms are offset by more than half a wavelength, causing the optimization to match the wrong cycle) is one of the central challenges in practical FWI.</p>

<p>A significant research effort over the past decade has gone into alternative misfit functions that are more convex than the $\ell^2$ norm: optimal transport distances (Wasserstein metrics), envelope-based misfits, and waveform-specific distance measures that are less sensitive to phase errors. The choice of misfit function turns out to matter enormously for whether gradient descent converges to a useful solution.</p>

<p>Ill-posedness is also a serious issue. With finite bandwidth, limited acquisition geometry, noise, and incomplete illumination, different velocity models can produce very similar recorded seismograms. Regularization is therefore essential, and the choice of regularizer encodes geological priors: smoothness in the background velocity, sharpness at interfaces, or more structured priors for salt geometries.</p>

<h2 id="where-this-is-used">Where This Is Used</h2>

<p><strong>Gulf of Mexico.</strong> Deepwater exploration in the Gulf is dominated by complex salt geology. Large salt domes with irregular geometries create strong velocity contrasts and shadow zones that conventional migration cannot image correctly. FWI with accurate salt flooding and velocity model building is the standard approach for imaging beneath and around these structures. Companies including Shell, BP, and Chevron run industrial-scale FWI workflows on survey areas covering thousands of square kilometers.</p>

<p><strong>Brazil’s pre-salt.</strong> The Santos and Campos basins offshore Brazil contain some of the world’s largest deepwater discoveries, located beneath thick layers of evaporites (salt) that make imaging extremely challenging. Petrobras and its partners have invested heavily in FWI-based velocity model building to map these reservoirs. The pre-salt fields are among the most seismically complex environments in which FWI has been applied at scale.</p>

<p><strong>Canada and the Arctic.</strong> The Canadian oil sands and the Beaufort Sea shelf represent different FWI challenges: land seismic in permafrost environments, and shallow-water marine seismic with near-surface velocity complexities. The Geological Survey of Canada and academic groups at the University of British Columbia have made significant contributions to FWI methodology for these settings.</p>

<p><strong>European North Sea.</strong> The Norwegian Continental Shelf (managed by Equinor and partners) is one of the most intensively surveyed and studied geological provinces in the world. Norwegian geophysics companies and university research groups (at the University of Bergen, NTNU, and the University of Oslo) have contributed substantially to FWI theory and practice, with a particular focus on elastic FWI (which accounts for both compressional and shear wave propagation) and anisotropy.</p>

<h2 id="the-machine-learning-connection">The Machine Learning Connection</h2>

<p>Classical FWI is iterative and gradient-based. The emerging research direction combines it with neural networks in several ways:</p>

<p><strong>Neural network velocity parametrization.</strong> Instead of optimizing $v(x)$ on a grid, parametrize it as the output of a neural network: $v(x) = \mathcal{N}_\theta(x)$. The optimization then runs over the network weights $\theta$. The implicit regularization of neural networks (their bias toward smooth functions) can act as a natural regularizer, and the network can be pretrained on geological databases to encode prior information.</p>

<p><strong>Physics-informed neural networks for forward modeling.</strong> PINNs can approximate solutions of wave equations and provide differentiable surrogates for inversion. However, high-frequency and multiscale wavefields remain difficult to train accurately, so PINNs do not yet replace established finite-difference, finite-element, or spectral solvers for large-scale FWI.</p>

<p><strong>Deep learning-based inversion.</strong> Train a neural network directly on pairs (seismograms, velocity models) to learn the inverse mapping. This requires large training datasets but can be extremely fast at inference time. The challenge is generalization: a network trained on a particular geological style may fail on a different basin.</p>

<p>Each of these directions is active, and none has displaced classical FWI in production workflows. Hybrid methods that combine adjoint-based optimization with learned components are a focused research direction because they retain the governing physics while using data to improve initialization, parametrization, or regularization.</p>

<p><em>For code related to seismic inversion and FWI, see the <a href="https://github.com/abdgafartunde/seismic-inversion-ml">seismic-inversion-ml GitHub repository</a> (to be updated).</em></p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Inverse Problems" /><category term="Geophysics" /><category term="Scientific Machine Learning" /><category term="Applied Mathematics" /><summary type="html"><![CDATA[How a PDE-constrained optimization problem sits at the heart of modern seismic exploration; solving it is one of the most computationally demanding inverse problems in industry.]]></summary></entry><entry><title type="html">Mathematics After Proof Scarcity: What Terence Tao’s ICM 2026 Lecture Means for Mathematicians</title><link href="https://abdgafartunde.github.io/blog/2026/08/17/mathematics-after-proof-scarcity/" rel="alternate" type="text/html" title="Mathematics After Proof Scarcity: What Terence Tao’s ICM 2026 Lecture Means for Mathematicians" /><published>2026-08-17T00:00:00+08:00</published><updated>2026-08-17T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/08/17/mathematics-after-proof-scarcity</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/08/17/mathematics-after-proof-scarcity/"><![CDATA[<p>Terence Tao gave a public lecture at the International Congress of Mathematicians 2026 titled <em>Mathematics in the Age of AI</em>. I expected a talk about how good the latest AI systems have become at mathematics: theorem proving, formalization, benchmarks, perhaps some predictions about where the technology is heading.</p>

<p>That is not really the talk he gave. Tao’s more interesting question was what happens <em>after</em> we grant that AI may become genuinely useful at research-level mathematics. Suppose, as a working hypothesis, that AI systems can soon perform a nontrivial fraction of mathematical research tasks with reasonable success, cost, and human supervision. What should mathematicians do then?</p>

<p>The question sounds practical, but it quickly becomes philosophical:</p>

<blockquote>
  <p><strong>What are the actual goals and values of mathematical research?</strong></p>
</blockquote>

<p>I have been thinking about this question since reading the slides, partly because it intersects with something that has bothered me for a while. I use AI in my own work. It helps me search literature, write and debug code, test ideas, check calculations, improve exposition, and sometimes explore mathematical arguments. The productivity gains are real. But there is an uncomfortable distinction between <em>producing mathematics</em> and <em>understanding mathematics</em>. We have historically been able to blur that distinction because producing a correct new result was difficult enough that the two often travelled together. AI may force us to separate them. This post is my attempt to work through Tao’s argument and what I think it means for researchers, particularly those of us working in applied mathematics, computation, inverse problems, and scientific machine learning.</p>

<h2 id="the-question-is-no-longer-only-whether-ai-can-do-mathematics">The Question Is No Longer Only Whether AI Can Do Mathematics</h2>

<p>A large fraction of the current debate is about capability. Can a language model really reason? Was a benchmark contaminated? Did the system solve a genuinely new problem? How much prompting was involved? Was the proof checked by a human or by a proof assistant? These are important questions. Tao formulates the issue as a family of “AI capability conjectures,” each a template whose key terms are left as placeholders: at some point, some AI tools will, at some expense, and with some level of human supervision, be able to accomplish some research-level mathematical tasks in some fields, with some non-trivial success rate and at some level of correctness. There is not one claim called “AI can do mathematics”; there are many, depending on which placeholders you fill in.</p>

<p>That distinction matters. Solving a carefully selected research problem after extensive human guidance is very different from autonomously developing a new mathematical theory. But Tao deliberately puts that dispute to one side. His argument is conditional. Suppose a reasonably strong version of the capability claim turns out to be true, and he offers one controlled data point: a second batch of ten novel research-level problems, assessed under scientific conditions against four AI systems in May 2026, with seven of the ten solved at publication-level quality, at compute costs of $10–$1000 per problem. Conditioning on that working hypothesis exposes a question that remains important even if one is sceptical about the strongest claims made by AI companies. The fact that a machine <em>can</em> perform a task does not tell us whether the task should be delegated to it, how its output should be evaluated, or what responsibilities remain with the mathematician. Capability is one question; value is another.</p>

<h2 id="mathematics-has-never-had-only-one-objective">Mathematics Has Never Had Only One Objective</h2>

<p>What are we trying to achieve when we do mathematical research? The obvious answer is “solve problems,” but that is only part of it. We also want to build theories, develop reusable techniques, understand phenomena, train future mathematicians, connect different areas of knowledge, support applications, and sometimes create something that is beautiful simply because of the structure it reveals.</p>

<p>Historically, these objectives were often positively correlated. An important problem was solved, and the proof introduced a technique. The technique became part of a theory. The theory generated further questions. Students learned it. Other researchers used it. Eventually the result became part of the standard language of the field. Because these goals tended to move together, the mathematical community could use visible outputs — solved problems, theorems, papers — as rough proxies for deeper progress.</p>

<p>AI may break that correlation. We can imagine a system that produces a large number of correct results while contributing relatively little to theory building, mathematical taste, education, or collective understanding. In that situation, maximizing the number of solved problems is no longer the same thing as maximizing mathematical progress.</p>

<p>This is where Tao invokes Goodhart’s law:</p>

<blockquote>
  <p><strong>When a measure becomes a target, it ceases to be a good measure.</strong></p>
</blockquote>

<p>The number of solved problems can be a useful measure when it emerges naturally from serious mathematical work. Once it becomes an explicit optimization target, a sufficiently capable system may become extremely good at increasing the count without necessarily increasing the thing we actually care about. That distinction feels increasingly important in academia more generally. We already know what happens when publications, citations, grant income, or benchmark scores become targets rather than indicators. AI can amplify the same problem by making the production of measurable output much cheaper. The danger is not only that AI might produce false mathematics. A more subtle danger is that it produces enormous quantities of <em>valid but low-value mathematics</em>.</p>

<h2 id="a-proof-is-not-the-end-of-the-process">A Proof Is Not the End of the Process</h2>

<p>One of the strongest parts of Tao’s lecture is his gradual reconstruction of what it actually means to solve a mathematical problem.</p>

<p>The naive picture is</p>

\[\text{open problem}
\longrightarrow
\text{solution}.\]

<p>But a proposed solution might be wrong. So we need verification:</p>

\[\text{open problem}
\longrightarrow
\text{unverified solution}
\longrightarrow
\text{verified solution}.\]

<p>Even this is not enough. Suppose an AI system produces a 150-page proof, and a formal proof assistant verifies every logical step. The theorem is now correct in a strong formal sense. But nobody understands the argument. Has the mathematical problem really been solved in the sense that matters to the discipline?</p>

<p>Tao’s answer is effectively no. The proof still needs exposition: it must be reorganized so that mathematicians can see the main mechanism, the difficult steps, the role of the assumptions, and the relationship with what was known before. Then it needs community acceptance. Experts must read it, test it, compare it with the literature, and decide whether it is important and trustworthy. And even publication is not the final stage. The strongest results are eventually <em>digested</em>: their proofs are simplified, their essential ideas extracted, they are placed in a more general framework, and they enter textbooks, graduate courses, surveys, software, formal libraries, and the working vocabulary of the field.</p>

<p>A more realistic pipeline is therefore</p>

\[\text{problem}
\rightarrow
\text{proof generation}
\rightarrow
\text{verification}
\rightarrow
\text{exposition}
\rightarrow
\text{acceptance}
\rightarrow
\text{digestion}
\rightarrow
\text{canonical theory}.\]

<p>I find this picture much more useful than the common debate about whether an AI has “solved” a theorem. It asks a better question: <strong>at which stage of the mathematical knowledge pipeline has the machine actually contributed?</strong></p>

<h2 id="from-proof-scarcity-to-proof-abundance">From Proof Scarcity to Proof Abundance</h2>

<p>Tao’s phrase that stayed with me most is the transition from <strong>proof scarcity</strong> to <strong>proof abundance</strong>. Our current mathematical institutions were built in a world where generating a serious new proof was expensive. The difficulty of creating the proof acted as a natural filter. There were still too many papers to read, of course, but the production rate was limited by the amount of human mathematical labour available. Suppose AI removes much of that bottleneck. Then the rate of proof generation could increase much faster than the rates of verification, exposition, peer review, and mathematical digestion.</p>

<p>One can think of the research system as a sequence of queues. Let $\lambda_g$ be the rate at which candidate proofs are generated, while $\mu_v$, $\mu_e$, $\mu_r$, $\mu_c$ represent our effective capacities for verification, exposition, review, and canonicalization. If AI makes $\lambda_g \gg \mu_v$, then unverified proofs accumulate. If verification is also automated but exposition remains slow, the bottleneck simply moves: $\lambda_v \gg \mu_e$. If AI becomes good at exposition too, journals and referees may become the limiting stage. And if reviewing is partially automated, we eventually encounter what seems to me the hardest bottleneck of all: human attention. There is only so much mathematics that a research community can genuinely absorb.</p>

<p>Tao calls the resulting phenomenon <strong>proof indigestion</strong>, and the term captures the problem well. Producing more mathematical objects does not automatically increase the amount of mathematics that the community understands. In fact, beyond some point, abundance can make understanding harder. Important results compete with thousands of technically correct but less consequential ones. Researchers spend more time filtering. The literature becomes harder to navigate. Priority becomes more difficult to establish. Expert refereeing becomes an increasingly scarce resource. The bottleneck shifts from <em>generation</em> to <em>judgment</em>.</p>

<h2 id="correctness-is-not-understanding">Correctness Is Not Understanding</h2>

<p>This distinction matters even more to me than proof abundance itself. A proof can be correct without being understood. Suppose a proof assistant verifies</p>

\[\Gamma \vdash T,\]

<p>where $\Gamma$ contains the assumptions and $T$ is the theorem. This establishes something very important: the formal derivation is valid relative to the encoded assumptions and definitions. But it does not answer why $T$ is true, which assumptions in $\Gamma$ are actually doing the work, what the central mechanism of the proof is, where the difficult step is, what would fail if one assumption were weakened, whether there is a stronger theorem hiding behind the argument, whether the proof reveals a reusable idea, or how it connects to the existing theory. These are not secondary questions. They are often where the mathematics lives.</p>

<p>When I read a good proof, I am rarely trying to memorize the sequence of deductions. I am trying to compress it into a mental model. Maybe the key is compactness. Maybe there is a hidden coercivity estimate. Maybe the right variable makes a convex structure visible. Maybe the whole argument is really exploiting an invariant. Maybe an apparently analytic theorem is ultimately geometric. Understanding occurs when the long derivation can be reorganized around a relatively small number of structural ideas. AI may become extremely good at producing derivations; whether that automatically produces this kind of conceptual compression is a separate question.</p>

<h2 id="the-strange-importance-of-friction">The Strange Importance of Friction</h2>

<p>Tao makes another observation that connects closely with something I wrote earlier in <a href="/blog/2026/02/06/the-great-deskilling/"><em>On the Quiet Erosion of Deep Thinking</em></a>. AI-generated mathematical writing can be extremely polished: the grammar is clean, the notation is consistent, every transition appears smooth. That sounds entirely desirable, but Tao points out that mathematical exposition can be <em>too smooth</em>.</p>

<p>In a human-written proof, the places where the author struggled often leave traces. There may be an extra paragraph explaining a subtle point, an awkward but revealing decomposition, a warning about a tempting false argument, or an example inserted exactly where intuition becomes difficult. These irregularities tell the reader where to slow down. A heavily AI-polished argument can remove both bad friction and useful friction: routine algebra and the genuinely new idea may be presented with the same confidence and at the same pace. The result is easy to read line by line while being surprisingly difficult to learn from.</p>

<p>This is a subtle point. We should not romanticize bad writing: confusing notation and unnecessary complication do not create depth. But there is a difference between removing obstacles to understanding and removing the evidence of where understanding is required. The best exposition does not make everything look easy; it makes the structure of the difficulty visible.</p>

<h2 id="this-is-also-a-training-problem">This Is Also a Training Problem</h2>

<p>The same issue appears in mathematical education. How did most of us learn to think mathematically? Not by continuously reading perfect solutions. We tried things that failed. We chose the wrong estimate. We constructed an example and discovered that our conjecture was false. We spent several days misunderstanding a definition before seeing why it had been formulated that way. Over time, those failures became judgment. The process looked roughly like</p>

\[\text{attempt}
\rightarrow
\text{failure}
\rightarrow
\text{diagnosis}
\rightarrow
\text{reformulation}
\rightarrow
\text{insight}.\]

<p>AI can intervene at every stage. That can be enormously helpful: a well-timed hint can save a student from wasting three days on a purely technical obstruction, a counterexample generated quickly can expose a false conjecture, and a different explanation can make an opaque definition understandable. But if the intervention occurs too early, the whole process collapses into</p>

\[\text{problem}
\rightarrow
\text{answer}.\]

<p>Then the student receives the result without developing the machinery that would have produced it. The interesting question is therefore not “Should mathematicians use AI?” (that question is already becoming outdated). The more important question is:</p>

<blockquote>
  <p><strong>Which cognitive operations must remain ours if we want to retain mathematical independence?</strong></p>
</blockquote>

<p>I do not yet have a complete answer. But I suspect the list includes problem formulation, recognizing structure, deciding which assumptions matter, developing examples, testing plausibility, detecting failure modes, and learning to remain productively stuck. These are not merely steps toward an answer; they are part of how mathematical taste is formed.</p>

<h2 id="authorship-has-to-mean-more-than-prompting">Authorship Has to Mean More Than Prompting</h2>

<p>Tao’s discussion of authorship is appropriately demanding. If AI contributes substantially to a mathematical result, what makes the human researcher the author? It cannot simply be that the researcher entered the prompt. Authorship must involve responsibility. A human author should be able to state the result precisely, explain the central mechanism, justify the assumptions, situate the work in the literature, identify what is genuinely new, answer expert questions, and correct the argument when a problem is discovered.</p>

<p>Tao offers a useful rule of thumb: if the authors cannot convincingly give a clear, correct, properly attributed expert-level talk on their own result, then the result should not be published under their names. I think this is a strong standard and the right direction. It shifts the criterion from “Did you personally type every line?” to something more substantive: <strong>do you possess the mathematics well enough to take responsibility for it?</strong> This matters especially with proprietary AI systems. A model may generate an argument without revealing where the idea came from, what related material it has seen, or whether part of the proof is effectively a rediscovery of something already in the literature. Disclosure of AI use is therefore necessary, but disclosure alone is not enough. We still have to reconstruct provenance, verify novelty, and understand the argument ourselves.</p>

<h2 id="the-problem-is-even-harder-in-applied-mathematics">The Problem Is Even Harder in Applied Mathematics</h2>

<p>In applied and computational mathematics, a formally correct theorem is only one layer of the problem.</p>

<p>Consider an inverse problem</p>

\[F(x)=y,\]

<p>where $F$ is the forward operator, $x$ is an unknown parameter or field, and $y$ is observed data.</p>

<p>With noisy measurements,</p>

\[\lVert y^\delta-y\rVert_Y\leq\delta,\]

<p>we might reconstruct $x$ by solving</p>

\[x_{\alpha,\theta}^{\delta}
\in
\operatorname*{arg\,min}_{x\in X}
\left\{
\mathcal{D}\bigl(F(x),y^\delta\bigr)
+
\alpha \mathcal{R}_{\theta}(x)
\right\}.\]

<p>An AI system could help derive this method, implement it, run the experiments, and perhaps even prove a convergence theorem. But none of that automatically answers the scientific questions: Is $F$ a sufficiently accurate model of the physical experiment? Is $\mathcal{D}$ the right model for the noise? What prior information is encoded in $\mathcal{R}_{\theta}$? Is the reconstruction identifiable from the available data? What happens under model mismatch? Does the discrete algorithm faithfully approximate the continuum formulation? Are the theoretical stability constants meaningful at computationally relevant scales? Does a learned regularizer still behave sensibly outside the training distribution?</p>

<p>These questions require judgment that lies outside the local correctness of the proof. This is one reason I think applied mathematicians should be especially careful about confusing AI-generated mathematical fluency with scientific understanding. The machine can manipulate the model we give it; we remain responsible for deciding whether it is the right model.</p>

<h2 id="what-should-become-more-valuable">What Should Become More Valuable?</h2>

<p>If proof generation becomes cheaper, the activities that remain scarce should become more valuable. I suspect we will need to place greater weight on the following.</p>

<p><strong>Problem selection.</strong> Knowing which questions are worth spending time on may become more important than executing every technical step of the solution.</p>

<p><strong>Theory building.</strong> A collection of isolated theorems is not a theory. Someone has to identify the right concepts, representations, invariants, and abstractions that organize them.</p>

<p><strong>Exposition.</strong> Not merely making arguments readable, but showing where the ideas are.</p>

<p><strong>Verification and reviewing.</strong> The community may need far more expert checking while the current incentive structure continues to treat reviewing as secondary service work.</p>

<p><strong>Synthesis and canonicalization.</strong> Surveys, textbooks, formal libraries, computational libraries, and definitive treatments may become increasingly important as the volume of raw output grows.</p>

<p><strong>Teaching and mentoring.</strong> If AI can supply answers instantly, helping students learn how to <em>think</em> becomes more important, not less.</p>

<p><strong>Scientific judgment.</strong> In applied mathematics, deciding whether a result is meaningful for the underlying phenomenon remains indispensable.</p>

<p>There is an interesting reversal here. The activities that have historically received less prestige because they happen <em>after</em> theorem generation may become the main bottlenecks of mathematical progress. In a world of abundant output, curation is not administrative cleanup; it is intellectual work.</p>

<h2 id="how-i-want-to-use-ai-in-my-own-research">How I Want to Use AI in My Own Research</h2>

<p>After reading Tao’s lecture, I do not feel any less inclined to use AI; if anything, I think mathematicians should become much better at using these systems. But I want the division of labour to be deliberate. I am comfortable delegating more of the mechanical work: code boilerplate, syntax, routine symbolic manipulation, literature discovery, formatting, preliminary numerical experiments, and checks that can be independently verified. I am much less comfortable delegating the parts that determine what the work <em>means</em>.</p>

<p>Before I accept an AI-assisted result as part of my own research, I want to be able to answer:</p>

<ol>
  <li>What is the precise mathematical or scientific question?</li>
  <li>Why should the proposed result be true?</li>
  <li>What is the main mechanism of the argument?</li>
  <li>Which assumptions are essential?</li>
  <li>What are the limiting or failure cases?</li>
  <li>How does the result relate to existing work?</li>
  <li>Can I reproduce the central reasoning without the AI output in front of me?</li>
  <li>Can I explain it clearly at the board to someone who knows the field?</li>
  <li>Can I tell which parts are routine and which parts are genuinely new?</li>
  <li>If the tool disappeared tomorrow, could I continue the project?</li>
</ol>

<p>That final question may be the most useful test. If the answer is no, then I may possess an output, but I do not yet possess the mathematics.</p>

<h2 id="what-i-think-taos-lecture-is-really-about">What I Think Tao’s Lecture Is Really About</h2>

<p>The title is <em>Mathematics in the Age of AI</em>, but Tao frames the lecture explicitly as describing a crisis in the foundations of mathematical values and practices, not a crisis of capability. This is an important distinction. The foundational crisis of roughly 1900–1930 forced mathematicians to make the foundations of reasoning explicit and ended by producing a rigorous, standardized framework. The present turbulence is, he argues, something different: it forces us to make explicit the values that were easy to leave implicit when human mathematical labour was the limiting resource.</p>

<p>What counts as understanding? What makes someone an author? What is a proof for? Why do we train mathematicians? What kinds of work should receive prestige? What turns an isolated theorem into mathematical knowledge? These questions were always there; AI makes them harder to avoid. And the urgency is real: Tao argues that mathematicians have only a narrow window to define what the profession means before those definitions get made for them by technology companies and financial incentives.</p>

<p>From this diagnosis he draws three concrete recommendations. First, normalize the responsible disclosure of AI assistance (covert use, concealed to dodge peer criticism, is the case to prevent, not AI use itself). Second, shift prestige away from proof generation and being “first,” and toward the slower human stages of exposition, publication, and canonicalization. Third, establish a publication gate: if authors cannot convincingly present their result at expert level, answer questions about it, and demonstrate command of the argument, the result should not be published under their names.</p>

<p>I do not think the answer is that mathematicians should compete with machines at symbolic speed, nor that we should retreat from AI and preserve an artificial version of twentieth-century mathematical practice. The more promising path is to use these systems aggressively where they extend our capabilities, while protecting the forms of reasoning, judgment, responsibility, and education on which meaningful mathematics depends. The future mathematician may perform fewer routine deductive steps manually; that does not necessarily make the mathematician less important. It may make the distinctly human parts of the job easier to see.</p>

<p>Choosing the right problem. Knowing when an answer is meaningful. Finding the idea inside a proof. Connecting isolated results into theory. Explaining why something matters. Teaching another person how to see it. Taking responsibility when the mathematics is wrong. Those are not peripheral activities surrounding theorem proving. They are part of what turns proofs into mathematics, and if Tao is right that we are moving from proof scarcity to proof abundance, they may become the most important parts of the profession.</p>

<hr />

<p><em>This essay is based on Terence Tao’s public lecture, “Mathematics in the Age of AI,” delivered at the International Congress of Mathematicians 2026 on 24 July 2026. The interpretation and reflections here are my own.</em></p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Mathematics" /><category term="Artificial Intelligence" /><category term="Research" /><category term="Reflections" /><summary type="html"><![CDATA[Reflections on Terence Tao's ICM 2026 lecture on mathematics in the age of AI: proof abundance, mathematical understanding, authorship, and what researchers should preserve as AI becomes more capable.]]></summary></entry><entry><title type="html">Reading Tikhonov (1963): The Paper That Made Ill-Posed Problems Solvable</title><link href="https://abdgafartunde.github.io/blog/2026/08/10/reading-tikhonov-1963/" rel="alternate" type="text/html" title="Reading Tikhonov (1963): The Paper That Made Ill-Posed Problems Solvable" /><published>2026-08-10T00:00:00+08:00</published><updated>2026-08-10T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/08/10/reading-tikhonov-1963</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/08/10/reading-tikhonov-1963/"><![CDATA[<p>Every field has a paper that changed the terms of the conversation. For the theory of ill-posed problems, that paper is Andrei Nikolaevich Tikhonov’s 1963 note in <em>Doklady Akademii Nauk SSSR</em>, translated into English as “On the regularization of ill-posed problems.” It is three and a half pages long. It introduced an idea that now bears his name, solved a problem that had been considered essentially intractable, and established the framework within which an entire subfield still operates.</p>

<p>I have cited this paper scores of times. When I finally sat down to read it carefully (not skim it, not rely on the secondary literature’s description of it, but actually read it), I was surprised by several things: what it said precisely, what it did not say, and how much conceptual content fits in so few pages. This post is a record of that close reading.</p>

<h2 id="the-problem-tikhonov-was-solving">The Problem Tikhonov Was Solving</h2>

<p>By 1963, the difficulty of ill-posed problems was well understood. Hadamard had articulated the concept of well-posedness in 1902 and used it, notoriously, to argue that ill-posed problems were not physically meaningful. His view was that a legitimate physical problem must have a solution that depends continuously on the data; problems that fail this continuity requirement were, in his view, wrongly formulated.</p>

<p>This position had become a genuine obstruction. Many problems of practical scientific interest (particularly problems of recovering internal structure from surface measurements) are ill-posed in Hadamard’s sense. The observed data determine the solution uniquely (in principle), but small errors in the data can produce arbitrarily large errors in the solution. If Hadamard was right that ill-posed problems were illegitimate, then these problems were unsolvable by design.</p>

<p>Tikhonov’s response was not to argue that Hadamard was wrong but to reframe the problem. His starting point: if the set of admissible solutions is constrained to a compact subset of the function space, then the inverse mapping is automatically continuous. The solution does depend stably on the data, but only among solutions that satisfy the constraint.</p>

<p>This is the conceptual core of regularization, stated in the very first paragraph of the 1963 paper.</p>

<h2 id="what-the-1963-paper-actually-contains">What the 1963 Paper Actually Contains</h2>

<p>The paper is short enough to describe precisely. Let me go through it.</p>

<p><strong>The setting.</strong> Tikhonov formulates the problem using an operator equation $Az = u$ between metric spaces. In the linear Banach-space specialization familiar today, $z$ is the unknown, $u$ is the data, and ill-posedness means that the inverse is not continuous on its range. Compact linear operators on infinite-dimensional spaces provide the standard example, but compactness of $A$ is not the general starting assumption.</p>

<p>He introduces a <strong>stabilizing functional</strong> $\Omega[z]$, a lower semicontinuous functional on the domain of $A$ whose sublevel sets ${z : \Omega[z] \leq M}$ are compact. These compact sets are the “admissible” solutions. The constraint $\Omega[z] \leq M$ encodes prior knowledge: you believe the true solution is regular in whatever sense $\Omega$ measures.</p>

<p><strong>The key construction.</strong> If $A$ is continuous and one-to-one on a compact admissible set, its inverse on the image of that set is continuous. Tikhonov then studies a regularizing functional that balances data discrepancy and the stabilizing functional. In Hilbert-space notation, the familiar quadratic form is</p>

\[R[z, u^\delta] = \lVert Az - u^\delta \rVert^2 + \alpha \Omega[z]\]

<p>has a minimizer for each $\alpha &gt; 0$ and $u^\delta$. This minimizer $z_\alpha^\delta$ is the regularized solution.</p>

<p>He then states the convergence result: if $\alpha = \alpha(\delta)$ is chosen so that $\alpha(\delta) \to 0$ and $\delta^2 / \alpha(\delta) \to 0$ as the noise level $\delta \to 0$, then $z_\alpha^\delta \to z^\dagger$ in norm, where $z^\dagger$ is the true solution.</p>

<p><strong>The Hilbert-space quadratic specialization.</strong> If $A$ is a bounded linear operator between Hilbert spaces and $\Omega[z] = \lVert Lz \rVert^2$, the first-order optimality condition has the familiar form</p>

\[(\alpha L^*L + A^*A) z = A^* u^\delta.\]

<p>This is the formula that now appears in every textbook on inverse problems. It is the equation whose solution produces the Tikhonov regularized estimate.</p>

<p><strong>An example.</strong> The paper closes with an example involving a first-kind integral equation. It illustrates how the stabilizing functional selects a stable approximate solution from noisy data.</p>

<h2 id="what-is-subtle-about-the-result">What Is Subtle About the Result</h2>

<p>Reading the paper carefully, a few things stand out that are sometimes obscured in textbook treatments.</p>

<p><strong>The parameter $\alpha$ is not determined by the paper.</strong> Tikhonov proves that <em>some</em> parameter choice strategy works (any $\alpha(\delta)$ with $\alpha \to 0$ and $\delta^2/\alpha \to 0$), but the paper gives no guidance on how to choose $\alpha$ in practice. The parameter choice problem (given actual noisy data $u^\delta$, how do you choose $\alpha$?) is entirely separate from the existence and convergence results. Methods for this (Morozov discrepancy principle, L-curve, generalized cross-validation) came later and are still active areas of research.</p>

<p><strong>The compact sublevel set assumption is strong.</strong> The convergence analysis relies on the sublevel sets of $\Omega$ being compact. For the $L^2$ norm in an infinite-dimensional Hilbert space, ${z : \lVert z \rVert^2 \leq M}$ is the closed ball, which is <em>not</em> compact. Tikhonov’s proof requires choosing $\Omega$ to be something like a Sobolev norm, whose sublevel sets are compact by the Rellich-Kondrachov theorem. The paper states this condition abstractly, but later textbooks sometimes state “Tikhonov regularization with $\Omega[z] = \lVert z \rVert^2$” without noting that the basic compactness argument no longer applies. The result is still true, but for different reasons.</p>

<p><strong>The word “regularization” is not in the title as a method, but as a process.</strong> Tikhonov is not proposing “a regularization method”; he is proving that ill-posed problems can be regularized (made well-posed) by a particular construction. The shift in meaning (from “regularization” as a property to “Tikhonov regularization” as a specific algorithm) happened gradually in the subsequent literature.</p>

<h2 id="what-the-paper-did-not-say">What the Paper Did Not Say</h2>

<p>The 1963 paper does not discuss:</p>

<ul>
  <li>
    <p><strong>Convergence rates.</strong> The paper proves that $z_\alpha^\delta \to z^\dagger$ as $\delta \to 0$ but does not quantify how fast. For standard quadratic Tikhonov regularization, the source condition $z^\dagger \in \mathcal{R}((A^*A)^\nu)$ gives the rate $\lVert z_\alpha^\delta - z^\dagger \rVert = O(\delta^{2\nu/(2\nu+1)})$ for $0 &lt; \nu \leq 1$, with an appropriate parameter choice. Later work developed these rate results and their extensions.</p>
  </li>
  <li>
    <p><strong>Nonlinear problems.</strong> The 1963 paper treats linear operators. Extending regularization theory to nonlinear operators is a much harder problem, and the general theory (convergence, rates, parameter choice for nonlinear Tikhonov) was developed primarily in the 1990s.</p>
  </li>
  <li>
    <p><strong>Discrepancy-based parameter choice.</strong> Morozov’s discrepancy principle, which chooses $\alpha$ so that $\lVert Az_\alpha^\delta - u^\delta \rVert \approx \delta$, was published by Morozov in 1966, three years later.</p>
  </li>
  <li>
    <p><strong>Statistical interpretation.</strong> The connection between Tikhonov regularization and MAP estimation under a Gaussian prior, which I described in earlier posts, was understood later. Tikhonov’s framework is deterministic.</p>
  </li>
</ul>

<h2 id="the-contemporaneous-soviet-context">The Contemporaneous Soviet Context</h2>

<p>One aspect of the paper that is easy to miss from a modern Western perspective is its place in Soviet mathematics.</p>

<p>Tikhonov was not primarily an inverse problems theorist; he was a mathematical physicist who had made fundamental contributions to topology, differential equations, and computational mathematics. His work on ill-posed problems was driven by real computational problems: the inversion of gravitational and electromagnetic data in geophysics, problems that the Soviet scientific establishment had strong incentives to solve.</p>

<p>The Soviet school of ill-posed problems (which included Tikhonov, Ivanov, Lavrentiev, and their students) developed in parallel with and largely independently of Western numerical analysis. The result was a body of theory that approached the same problems from a different angle, emphasizing functional analysis and operator theory rather than matrix computation. The convergence of the two traditions in the 1980s and 1990s enriched both.</p>

<p>Tikhonov’s 1963 paper was published in Russian; the English translation appeared in <em>Soviet Mathematics Doklady</em> the same year. The rapid availability of translations was partly a feature of the Cold War scientific communication apparatus, which made Soviet mathematics relatively accessible to Western readers despite the political context.</p>

<h2 id="why-the-paper-still-matters">Why the Paper Still Matters</h2>

<p>By the standards of what we know now, the 1963 paper is incomplete. It addresses the linear case, gives no convergence rates, offers no practical parameter choice method, and works under a compactness assumption that modern analyses often bypass. The theory has been extended in every direction.</p>

<p>And yet the paper is still worth reading, for two reasons.</p>

<p>First, the core insight is elegant and clear: compactness restores stability. Once you have seen this, the whole framework of regularization makes sense. You are not applying an algorithm; you are enforcing a constraint that makes the problem well-posed. Every subsequent development (sparsity constraints, total variation, machine-learned regularizers) is a different way of choosing what that constraint should be.</p>

<p>Second, the paper is a model of how to introduce a new framework. Tikhonov states the problem clearly, identifies the key obstruction (unboundedness of the inverse), introduces the minimal modification needed to restore stability, proves the fundamental theorem, and gives an example. There is no unnecessary generality, no anticipation of future extensions, no hedging. The result fills three and a half pages and answers the question it set out to answer.</p>

<p>That clarity is harder to achieve than it appears, and it is worth studying independently of the mathematics.</p>

<h2 id="further-context">Further Context</h2>

<p>The bibliographic record and Russian full text of Tikhonov’s note are available from <a href="https://www.mathnet.ru/eng/dan/v153/i1/p49">Math-Net.Ru</a>. For a modern treatment of regularization theory, Engl, Hanke, and Neubauer’s <em>Regularization of Inverse Problems</em> (1996) is a standard reference. For the Bayesian counterpart, Stuart’s 2010 <em>Acta Numerica</em> paper <a href="https://doi.org/10.1017/S0962492910000061">“Inverse problems: a Bayesian perspective”</a> is a useful entry point.</p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Mathematics" /><category term="Inverse Problems" /><category term="History of Mathematics" /><summary type="html"><![CDATA[A close reading of Tikhonov's 1963 paper on the regularization of ill-posed problems: what it actually says, what it assumed, and why it changed how we think about inverse problems.]]></summary></entry><entry><title type="html">Stochastic PDEs: When the Equations Themselves Are Random</title><link href="https://abdgafartunde.github.io/blog/2026/07/27/stochastic-pdes/" rel="alternate" type="text/html" title="Stochastic PDEs: When the Equations Themselves Are Random" /><published>2026-07-27T00:00:00+08:00</published><updated>2026-07-27T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/07/27/stochastic-pdes</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/07/27/stochastic-pdes/"><![CDATA[<p>Most of the PDEs I work with are deterministic: given a fixed conductivity $\sigma$, the potential $u$ is the unique solution of an elliptic equation. But the real world introduces randomness at every level. The material properties are measured imprecisely. The geometry is known only approximately. The applied currents are subject to instrument noise. The boundary conditions depend on environmental conditions that fluctuate.</p>

<p>One response to this is what I normally do: treat the PDE deterministically and handle uncertainty through the inversion process: Bayesian inference, regularization, error bounds. But there is an alternative: incorporate the randomness directly into the PDE. The result is a <strong>stochastic PDE</strong> (SPDE), a PDE whose coefficients, forcing, or boundary conditions are random processes.</p>

<p>This framework is useful whenever a linearization around a single point estimate does not capture the effect of uncertain inputs, or whenever the statistical properties of the solution are themselves of interest. It also provides a mathematical model for physical phenomena such as reaction-diffusion in heterogeneous media and wave propagation in random media, where randomness is part of the model rather than only measurement error.</p>

<h2 id="two-flavours-of-stochastic-pdes">Two Flavours of Stochastic PDEs</h2>

<p>There is an important distinction between two types of randomness in PDEs.</p>

<p><strong>Parametric uncertainty (random coefficients).</strong> The PDE has deterministic structure but uncertain parameters:</p>

\[-\nabla \cdot (a(\omega, x) \nabla u(\omega, x)) = f(x) \quad \text{in } \Omega,\]

<p>where $a(\omega, x)$ is a random field indexed by $\omega \in \Omega_{\text{prob}}$ (the probability space). For each realization $\omega$, this is a standard deterministic PDE. The randomness is in the coefficient, not in the differential structure.</p>

<p>This is the setting for uncertainty quantification: you have a model, uncertain inputs, and you want to understand how the uncertainty propagates to the output. The solution $u(\omega, \cdot)$ is a random field, and you want its statistics (mean, variance, distribution of quantities of interest).</p>

<p><strong>Stochastic forcing (noise-driven PDEs).</strong> The PDE is driven by a random forcing term, typically modelled as (Gaussian) white noise or a spatially correlated noise:</p>

\[\frac{\partial u}{\partial t} = \mathcal{L} u + \dot{W}(t, x),\]

<p>where $\dot{W}$ denotes spacetime white noise, interpreted as a generalized Gaussian random field with formal covariance $\mathbb{E}[\dot{W}(t,x)\dot{W}(s,y)] = \delta(t-s)\delta(x-y)$. Depending on the equation and the spatial dimension, the solution may be a function-valued stochastic process or only a distribution-valued one.</p>

<p>The two settings require different mathematical tools, though they share themes. I will focus primarily on the parametric setting, which is more directly connected to my work.</p>

<h2 id="random-fields-and-karhunen-loève-expansion">Random Fields and Karhunen-Loève Expansion</h2>

<p>A central object in the parametric SPDE setting is a <strong>random field</strong>: a function $a : \Omega_\text{prob} \times D \to \mathbb{R}$ that is random in $\omega$ and spatially distributed in $x \in D$.</p>

<p>For a Gaussian random field with mean $\bar{a}(x)$ and covariance function $C(x, x’) = \text{Cov}(a(\cdot, x), a(\cdot, x’))$, the <strong>Karhunen-Loève (KL) expansion</strong> provides a canonical series representation:</p>

\[a(\omega, x) = \bar{a}(x) + \sum_{k=1}^\infty \sqrt{\lambda_k} \xi_k(\omega) \psi_k(x),\]

<p>where $(\lambda_k, \psi_k)$ are the eigenvalue-eigenfunction pairs of the covariance operator</p>

\[(C\psi)(x) = \int_D C(x, x') \psi(x')\, dx',\]

<p>and $\xi_k(\omega) \sim \mathcal{N}(0, 1)$ are independent standard Gaussians. For a square-integrable random field with a trace-class covariance operator, the KL expansion converges in mean square. Mercer’s theorem gives the corresponding eigenfunction expansion when the covariance kernel is continuous and positive definite on a compact domain.</p>

<p>The KL expansion separates the randomness (the $\xi_k$) from the spatial structure (the $\psi_k$). If the covariance is smooth, the eigenvalues $\lambda_k$ decay rapidly, and the field is well-approximated by truncating the series at $K$ terms:</p>

\[a_K(\omega, x) = \bar{a}(x) + \sum_{k=1}^K \sqrt{\lambda_k} \xi_k(\omega) \psi_k(x).\]

<p>This truncation is the key step: it replaces the infinite-dimensional random field by a $K$-dimensional random vector $\boldsymbol{\xi} = (\xi_1, \ldots, \xi_K)$. The PDE solution $u(\omega, \cdot)$ becomes a function of $\boldsymbol{\xi}$, and the problem reduces to understanding the function $u(\boldsymbol{\xi}, \cdot)$ defined on $\mathbb{R}^K$.</p>

<h2 id="uncertainty-quantification-forward-and-inverse">Uncertainty Quantification: Forward and Inverse</h2>

<p><strong>Forward UQ</strong> propagates input uncertainty to output uncertainty. Given the random coefficient $a(\omega, x)$, what is the distribution of the solution $u(\omega, x)$, or of a quantity of interest $Q(\omega) = \mathcal{Q}(u(\omega))$?</p>

<p>The simplest approach is <strong>Monte Carlo</strong>: sample $\boldsymbol{\xi}^{(1)}, \ldots, \boldsymbol{\xi}^{(M)}$ from their distribution, solve the PDE for each sample, and average. The expected value is approximated by $\bar{Q} \approx \frac{1}{M}\sum_{i=1}^M Q(\boldsymbol{\xi}^{(i)})$. The root-mean-square sampling error decays as $O(M^{-1/2})$, while the mean-square error decays as $O(M^{-1})$, provided $Q$ has finite variance. These rates do not depend explicitly on the dimension $K$. Monte Carlo is easy to parallelize, but its sampling error decreases slowly.</p>

<p><strong>Stochastic Galerkin methods</strong> represent $u(\boldsymbol{\xi}, x)$ as a polynomial in $\boldsymbol{\xi}$:</p>

\[u(\boldsymbol{\xi}, x) \approx \sum_{|\alpha| \leq p} u_\alpha(x) \Psi_\alpha(\boldsymbol{\xi}),\]

<p>where $\Psi_\alpha$ are multivariate Hermite (or Legendre) polynomials and $\alpha$ is a multi-index. The coefficients $u_\alpha(x)$ are deterministic functions, solved for by a Galerkin projection of the SPDE. The resulting system couples all polynomial modes, producing a large but structured linear system. For smooth dependence of $u$ on $\boldsymbol{\xi}$, the polynomial approximation achieves spectral convergence in the stochastic dimension.</p>

<p><strong>Stochastic collocation</strong> evaluates the PDE at a set of quadrature points in $\boldsymbol{\xi}$-space, then interpolates or averages. Unlike Galerkin, it is non-intrusive: it reuses the deterministic PDE solver as a black box. Sparse-grid quadrature rules can reduce the cost when the solution depends smoothly and anisotropically on a moderate number of influential parameters. There is no universal dimension threshold; feasibility depends on regularity and effective dimension.</p>

<p><strong>Inverse UQ</strong> combines the forward UQ framework with Bayesian inference. The unknown coefficient $a(\omega, x)$ is treated as a random field to be inferred from data. The KL expansion parameterizes the unknown, the Bayesian posterior gives the distribution of $\boldsymbol{\xi}$ given the measurements, and MCMC or deterministic inference methods (Laplace approximation, transport maps) are used to characterize the posterior. This is the framework I described in the Bayesian inverse problems post, extended to the function-space setting.</p>

<h2 id="the-challenge-of-infinite-dimensions">The Challenge of Infinite Dimensions</h2>

<p>The SPDE setting raises mathematical issues that finite-dimensional stochastic analysis does not. Spacetime white noise $\dot{W}(t, x)$ is not a genuine function; it is only a distribution. Solutions to noise-driven PDEs may not be pointwise-defined functions but only elements of suitable Sobolev spaces.</p>

<p>For the stochastic heat equation $\partial_t u = \Delta u + \dot{W}$ in dimension $d$, the solution is a pointwise-defined function for $d = 1$, a distribution for $d \geq 2$. In the stochastic wave equation and nonlinear SPDEs, the situation is more delicate.</p>

<p>This affects how numerical results are interpreted. A discretization of spacetime white noise is mesh-dependent, and convergence must be understood in an appropriate function or distribution space. Linear equations with additive noise can often be treated through mild or weak solutions. Singular nonlinear SPDEs are harder because products of distributions may be undefined; renormalization techniques such as Wick products, regularity structures, or paracontrolled distributions are then required.</p>

<p>The development of a theory of singular SPDEs, including Martin Hairer’s theory of regularity structures, resolved long-standing questions about nonlinear equations driven by rough noise. Hairer received the Fields Medal in 2014 in part for this work. The theory is technically demanding and has reshaped the modern analysis of singular stochastic equations.</p>

<h2 id="connection-to-gaussian-process-priors">Connection to Gaussian Process Priors</h2>

<p>There is a precise connection between SPDEs and the Gaussian process priors I described in an earlier post. Lindgren, Rue, and Lindström (2011) showed that a Gaussian Markov random field (GMRF) defined by the SPDE</p>

\[(\kappa^2 - \Delta)^{\alpha/2} u = \mathcal{W}\]

<p>has a Matérn covariance with smoothness parameter $\nu = \alpha - d/2$, where $d$ is the spatial dimension and $\mathcal{W}$ is Gaussian white noise. This connection has practical consequences: the SPDE can be discretized using FEM to produce a sparse precision matrix (the inverse of the covariance matrix), which allows Gaussian process inference to scale to spatial problems with millions of observation locations.</p>

<p>For prior construction in inverse problems, this means that a Matérn GP prior can be implemented without forming or inverting the dense covariance matrix: instead, you work with the sparse FEM precision matrix directly. This is orders of magnitude cheaper and handles irregular geometries naturally.</p>

<h2 id="what-this-framework-offers">What This Framework Offers</h2>

<p>The stochastic PDE framework offers something that standard deterministic inverse problems, by themselves, do not: a principled account of what it means to have uncertain model parameters, not just uncertain data. When the conductivity in EIT is genuinely variable (biological tissue has natural variability), treating it as fixed but unknown and recovering a single best estimate is a modelling choice that can be questioned. The stochastic framework takes a different position: the conductivity is itself a random field, and the goal is to characterize its distribution.</p>

<p>Whether this framework is appropriate depends on the application. For quality control (each manufactured part has a fixed conductivity, I want to find it), deterministic inversion is correct. For biological imaging (each patient’s tissue properties are drawn from a population distribution), the stochastic framework is more natural. The right choice depends on what question you are actually trying to answer, and being explicit about that question is, I have found, one of the most clarifying exercises in any modelling problem.</p>

<hr />

<h2 id="stochastic-pdes-in-science-and-engineering">Stochastic PDEs in Science and Engineering</h2>

<p>Randomness in equations is not a limitation to be apologized for. In many of the most consequential scientific and engineering systems operating today, it is simply an accurate description of reality.</p>

<p><strong>Earthquake hazard in Japan.</strong> Seismic waves propagate through a heterogeneous crust whose density, wave speed, and attenuation are not known exactly. Probabilistic wave-propagation and uncertainty-quantification models can therefore complement deterministic simulations when estimating arrival times and shaking intensity. Operational early-warning systems also depend on rapid detection, empirical prediction models, and real-time sensor data, so they should not be described as direct SPDE solvers.</p>

<p><strong>Flood and typhoon modelling.</strong> Operational weather centres use ensemble prediction, integrating atmospheric models from perturbed initial states and sometimes with perturbed model parameters. These are often ensembles of deterministic PDE solves rather than direct discretizations of an SPDE. The ensemble spread is nevertheless an important estimate of forecast uncertainty for rainfall, river flow, and tropical-cyclone tracks. Stochastic parameterizations provide one way to represent unresolved processes within such models.</p>

<p><strong>Semiconductor manufacturing.</strong> Depositing thin films and etching features at nanometre scales involves uncertain reaction and transport processes. Stochastic continuum and particle models are used in research to study how fluctuations in feature size and position affect manufacturing yield. Claims about the internal process-control models used by particular companies require public technical sources and should not be inferred from the research literature alone.</p>

<p><strong>Financial mathematics.</strong> The Black-Scholes pricing equation is a deterministic PDE derived from a stochastic differential equation for the underlying asset. Heston and SABR are stochastic differential equation models, not SPDEs. SPDEs arise in more elaborate models where the evolving state is itself a function, for example a stochastic forward-rate curve.</p>

<p>The stochastic PDE framework provides the mathematical language for any system where the governing equations are themselves subject to uncertainty, whether that uncertainty arises from incomplete knowledge of material properties, chaotic dynamics, external forcing, or intrinsic randomness at the molecular scale.</p>

<p>For the link between Matérn fields and sparse precision operators, see Lindgren, Rue, and Lindström, <a href="https://doi.org/10.1111/j.1467-9868.2011.00777.x">“An explicit link between Gaussian fields and Gaussian Markov random fields: the stochastic partial differential equation approach”</a> (2011).</p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Mathematics" /><category term="Stochastic PDEs" /><category term="Uncertainty Quantification" /><summary type="html"><![CDATA[An introduction to PDEs with random inputs: why randomness arises, how solutions are defined, and the connections to uncertainty quantification and inverse problems.]]></summary></entry><entry><title type="html">Spectral Methods: Exponential Accuracy from Smooth Solutions</title><link href="https://abdgafartunde.github.io/blog/2026/07/13/spectral-methods/" rel="alternate" type="text/html" title="Spectral Methods: Exponential Accuracy from Smooth Solutions" /><published>2026-07-13T00:00:00+08:00</published><updated>2026-07-13T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/07/13/spectral-methods</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/07/13/spectral-methods/"><![CDATA[<p>In my earlier post on the finite element method, I described approximation by piecewise polynomials. For conforming degree-$p$ elements applied to a sufficiently regular elliptic solution on a shape-regular mesh, the energy-norm error is typically $O(h^p)$. Other norms and reduced regularity give different rates.</p>

<p>This is algebraic convergence under fixed polynomial degree. On a quasi-uniform mesh in $d$ dimensions, replacing $h$ by $h/2^8$ increases the number of degrees of freedom by approximately $2^{8d}$ and reduces an $O(h^p)$ error by $2^{8p}$. Adaptive meshes and higher-order finite elements change this comparison.</p>

<p>Spectral methods offer a different convergence regime. Analytic solutions can yield exponential coefficient decay and errors of the form $O(e^{-cN})$ in one-dimensional truncation order $N$. Finitely differentiable solutions give algebraic rates, while infinitely differentiable nonanalytic solutions generally give superalgebraic rather than necessarily exponential convergence. The comparison with finite elements depends on dimension, regularity, geometry, and whether high-order or spectral elements are used.</p>

<h2 id="the-core-idea">The Core Idea</h2>

<p>A spectral method approximates the solution $u$ of a PDE by a linear combination of globally supported basis functions:</p>

\[u_N(x) = \sum_{k=0}^{N} \hat{u}_k \phi_k(x),\]

<p>where ${\phi_k}$ is a sequence of smooth, globally defined functions (Fourier modes, Chebyshev polynomials, Legendre polynomials), and $\hat{u}_k$ are the expansion coefficients to be determined.</p>

<p>The contrast with finite elements is immediate: FEM uses piecewise polynomial basis functions that are zero outside a small patch of elements. Spectral basis functions are nonzero everywhere in the domain. Each spectral basis function “sees” the entire solution, which is why spectral methods can extract more information from a given number of degrees of freedom.</p>

<h2 id="fourier-methods-on-periodic-domains">Fourier Methods on Periodic Domains</h2>

<p>On the interval $[0, 2\pi]$ with periodic boundary conditions, the natural spectral basis is the Fourier series:</p>

\[u_N(x) = \sum_{k=-N/2}^{N/2} \hat{u}_k e^{ikx}.\]

<p>The coefficients $\hat{u}_k$ are the Fourier coefficients of $u$. The quality of the approximation is controlled by how fast these coefficients decay. For a function with $p$ continuous derivatives, $\lvert\hat{u}_k\rvert = O(\lvert k\rvert^{-p})$ as $\lvert k\rvert \to \infty$, giving algebraic decay. For a real analytic function (one that extends to a holomorphic function in a strip around the real axis), the coefficients decay exponentially: $\lvert\hat{u}_k\rvert \leq C e^{-\sigma \lvert k\rvert}$ for some $\sigma &gt; 0$. This exponential decay of coefficients is the origin of spectral convergence.</p>

<p><strong>Differentiation.</strong> For a periodic Fourier series, the $k$-th coefficient of $u’$ is $ik\hat{u}_k$. This identity is exact for the truncated representation. Truncation, aliasing, roundoff, and the conditioning of discrete differentiation still affect a numerical method.</p>

<p><strong>The FFT.</strong> Evaluating a Fourier series at $N$ equally spaced points, and computing the Fourier coefficients from $N$ function values, costs $O(N \log N)$ via the Fast Fourier Transform. This makes Fourier spectral methods exceptionally efficient: the cost per degree of freedom is nearly optimal.</p>

<p><strong>Gibbs phenomenon.</strong> Near an isolated jump, Fourier partial sums exhibit a limiting overshoot of about 9% of the jump height. For a piecewise smooth function with a nonzero jump, the coefficients generally decay as $O(\lvert k\rvert^{-1})$. The convergence rate depends on the norm: the global $L^2$ truncation error is typically of order $N^{-1/2}$, while pointwise convergence fails at the jump and is nonuniform nearby.</p>

<h2 id="chebyshev-and-legendre-methods-on-bounded-intervals">Chebyshev and Legendre Methods on Bounded Intervals</h2>

<p>For non-periodic problems on $[-1, 1]$, the natural spectral bases are orthogonal polynomials. The two most widely used are:</p>

<p><strong>Chebyshev polynomials</strong> $T_k(x) = \cos(k \arccos x)$. They are defined by the recurrence $T_0 = 1$, $T_1 = x$, $T_{k+1}(x) = 2xT_k(x) - T_{k-1}(x)$. The Chebyshev nodes $x_j = \cos(\pi j / N)$ cluster near the endpoints $\pm 1$, which is essential for controlling the Runge phenomenon.</p>

<p><strong>Legendre polynomials</strong> $P_k(x)$, defined by the recurrence $P_0 = 1$, $P_1 = x$, $(k+1)P_{k+1} = (2k+1)xP_k - kP_{k-1}$. They are $L^2$-orthogonal with respect to the uniform weight on $[-1, 1]$.</p>

<p>Both families achieve spectral convergence for smooth functions: if $u$ is analytic on $[-1, 1]$, the expansion coefficients decay exponentially and the $N$-term approximation satisfies</p>

\[\lVert u - u_N \rVert_{L^\infty} \leq C e^{-\sigma N}\]

<p>for constants $C, \sigma &gt; 0$ depending on the analytic continuation of $u$.</p>

<p>The Chebyshev expansion has the practical advantage that it reduces to a cosine transform: evaluating the expansion at Chebyshev nodes costs $O(N \log N)$ via the DCT (discrete cosine transform). Chebyshev methods are also closely related to Fourier methods through the substitution $x = \cos\theta$.</p>

<h2 id="solving-pdes-the-galerkin-and-collocation-approaches">Solving PDEs: The Galerkin and Collocation Approaches</h2>

<p>Given a basis ${\phi_k}$, there are two standard ways to enforce the PDE.</p>

<p><strong>Galerkin spectral method.</strong> Require the residual to be orthogonal to the span of the first $N$ basis functions:</p>

\[\langle F(u_N), \phi_j \rangle = 0 \quad \text{for } j = 0, \ldots, N.\]

<p>For a linear PDE $Lu = f$, this gives the system $\langle L u_N, \phi_j \rangle = \langle f, \phi_j \rangle$, which in matrix form is a dense linear system for the coefficients $\hat{u}_k$. The resulting matrix is often well-conditioned for elliptic problems when the right basis is used.</p>

<p><strong>Spectral collocation (pseudospectral) method.</strong> Require the PDE to be satisfied exactly at a set of collocation points ${x_j}_{j=0}^N$:</p>

\[(Lu_N)(x_j) = f(x_j) \quad \text{for all } j.\]

<p>The collocation points are typically the Gauss quadrature nodes (roots of $\phi_{N+1}$) or the Gauss-Lobatto nodes (roots of $(1-x^2)\phi_N’$). Collocation is often simpler to implement than Galerkin for variable-coefficient or nonlinear problems because it avoids computing inner products analytically.</p>

<p>The two approaches are closely related and have similar convergence properties. In practice, spectral collocation is more widely used in fluid dynamics and wave equations; Galerkin is more common in problems where variational structure is important.</p>

<h2 id="where-spectral-methods-excel">Where Spectral Methods Excel</h2>

<p><strong>Regular solutions on simple geometries.</strong> Fourier methods fit periodic rectangular domains, while polynomial spectral methods fit intervals and tensor-product domains. Analytic data and solutions can produce exponential convergence, but the number of modes needed for a given tolerance depends on singularities in the complex plane, boundary conditions, conditioning, and dimension.</p>

<p><strong>Fluid dynamics.</strong> Spectral methods have dominated direct numerical simulation of turbulence since the 1970s. For the Navier-Stokes equations in a box with periodic boundary conditions, Fourier methods are the tool of choice. Many of the benchmark results in turbulence research were computed spectrally.</p>

<p><strong>Wave propagation.</strong> High-order accuracy matters for wave equations because low-order methods introduce numerical dispersion: the computed wave speed differs from the true wave speed, and the error accumulates over long propagation distances. Spectral methods, with their exponential accuracy for smooth solutions, suffer far less numerical dispersion.</p>

<h2 id="where-spectral-methods-struggle">Where Spectral Methods Struggle</h2>

<p><strong>Complex geometry.</strong> A Fourier method lives on a rectangle, a Chebyshev method on an interval. Handling L-shaped domains, holes, or patient-specific geometry requires either domain decomposition (spectral element methods) or coordinate transformations that can spoil the exponential convergence. Finite elements handle complex geometry naturally; spectral methods do not.</p>

<p><strong>Discontinuous solutions.</strong> As noted, Gibbs oscillations contaminate the solution near any discontinuity. Remedies exist (filtering, WENO reconstruction, Gegenbauer post-processing) but they add complexity and do not fully restore spectral convergence.</p>

<p><strong>Preconditioning.</strong> The dense matrices arising in spectral methods can be ill-conditioned, especially for high-order problems. Designing efficient preconditioners is more difficult than for FEM, where the sparse, banded structure facilitates many standard preconditioner constructions.</p>

<h2 id="why-i-know-this">Why I Know This</h2>

<p>Most of my direct numerical work uses finite element methods, because EIT involves complex boundary shapes and potentially discontinuous conductivities, exactly the settings where FEM excels over spectral methods. But spectral methods appear indirectly in my work in two ways.</p>

<p>First, Fourier approximation and spectral regularization both analyze expansions in basis or singular vectors, but their convergence theories are not identical. Spectral cutoff for an inverse problem is governed by singular-value decay, noise, and source conditions, whereas Fourier approximation is governed primarily by function regularity.</p>

<p>Second, the spectral element method (SEM), which combines spectral accuracy within each element with FEM flexibility in geometry, is increasingly used for forward solvers in wave-based inverse problems. Understanding spectral methods makes SEM more transparent.</p>

<p>Spectral methods make the role of regularity in approximation explicit. Finite smoothness gives algebraic decay, additional smoothness can give superalgebraic decay, and analyticity can give exponential decay. This relationship appears throughout approximation theory and operator equations.</p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Mathematics" /><category term="Scientific Computing" /><category term="Spectral Methods" /><summary type="html"><![CDATA[An introduction to spectral methods for PDEs: when they achieve rapid or exponential convergence, how they compare with finite elements, and where they struggle.]]></summary></entry><entry><title type="html">Optimal Experimental Design: Choosing Measurements That Matter</title><link href="https://abdgafartunde.github.io/blog/2026/06/29/optimal-experimental-design/" rel="alternate" type="text/html" title="Optimal Experimental Design: Choosing Measurements That Matter" /><published>2026-06-29T00:00:00+08:00</published><updated>2026-06-29T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/06/29/optimal-experimental-design</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/06/29/optimal-experimental-design/"><![CDATA[<p>Most discussions of inverse problems take the measurement setup as given. You have a set of electrodes, or a set of transducers, or a set of sensors, and you work with whatever data they produce. But there is a prior question: which measurements should you take in the first place?</p>

<p>This is the problem of <strong>optimal experimental design</strong> (OED). Given a model, a class of possible experiments, and a criterion for what “good” information looks like, OED asks: which experiment (which choice of measurement operator) extracts the most useful information about the unknown?</p>

<p>The question has a long history in statistics, going back at least to work by Kiefer and Wolfowitz in the 1950s. In recent years, it has become practically significant in inverse problems, where the measurement geometry is often a design choice, not a physical constraint. In EIT, you can choose which electrode pairs to stimulate and in what pattern. In seismic surveys, you can choose the shot locations. In MRI, you can choose the $k$-space trajectory. In each case, making the right choices can dramatically improve the quality of the reconstruction.</p>

<h2 id="the-setup">The Setup</h2>

<p>Suppose the unknown parameter $m$ takes values in $\mathbb{R}^n$, and we observe linear data</p>

\[y = Cm + \eta,\]

<p>where $C \in \mathbb{R}^{k \times n}$ is the measurement operator (determined by the experimental design), and $\eta \sim \mathcal{N}(0, \Gamma_\text{noise})$ is Gaussian noise.</p>

<p>We adopt a Bayesian viewpoint. The prior is $m \sim \mathcal{N}(m_0, \Gamma_\text{prior})$. Given data $y$, the posterior is also Gaussian:</p>

\[m \mid y \sim \mathcal{N}(m_\text{post}, \Gamma_\text{post}),\]

<p>with</p>

\[\Gamma_\text{post}^{-1} = \Gamma_\text{prior}^{-1} + C^\top \Gamma_\text{noise}^{-1} C.\]

<p>The posterior covariance $\Gamma_\text{post}$ is the key object. It measures the remaining uncertainty after the experiment. A good experiment drives $\Gamma_\text{post}$ toward zero in the directions we care about; a bad experiment leaves large posterior variance.</p>

<p>The problem of OED is to choose $C$ (subject to a budget constraint) to minimize a scalar summary of $\Gamma_\text{post}$.</p>

<h2 id="design-criteria">Design Criteria</h2>

<p>Several scalar criteria for $\Gamma_\text{post}$ are standard.</p>

<p><strong>A-optimality</strong> minimizes the trace of $\Gamma_\text{post}$, which is the average posterior variance across all components of $m$:</p>

\[\Phi_A(C) = \text{tr}(\Gamma_\text{post}) = \text{tr}\!\left(\left(\Gamma_\text{prior}^{-1} + C^\top \Gamma_\text{noise}^{-1} C\right)^{-1}\right).\]

<p>Under the assumed linear-Gaussian model and squared Euclidean loss, minimizing the trace minimizes the Bayes risk of the posterior mean.</p>

<p><strong>D-optimality</strong> minimizes the log-determinant of $\Gamma_\text{post}$, equivalently maximizing the information gain (the reduction in differential entropy from prior to posterior):</p>

\[\Phi_D(C) = \log\det(\Gamma_\text{post}) = -\log\det\!\left(\Gamma_\text{prior}^{-1} + C^\top \Gamma_\text{noise}^{-1} C\right).\]

<p>The information gain is $\frac{1}{2}\log\det(\Gamma_\text{prior}) - \frac{1}{2}\log\det(\Gamma_\text{post})$, which measures (in nats) how much entropy is removed by the experiment. D-optimal designs maximize this reduction, and they have a Bayesian-invariance property: the optimal design is the same regardless of the prior’s mean $m_0$.</p>

<p><strong>E-optimality</strong> minimizes the largest eigenvalue of $\Gamma_\text{post}$, reducing the worst-case posterior variance direction. This is more conservative than A-optimality and appropriate when the largest uncertainty matters more than the average.</p>

<p><strong>Goal-oriented design</strong> targets a specific quantity of interest $Q(m) = q^\top m$ and minimizes $q^\top \Gamma_\text{post} q$, the posterior variance of that quantity. This is appropriate when only part of the unknown is relevant; global design criteria can allocate measurement effort to directions that do not affect $Q$.</p>

<h2 id="the-sensor-placement-problem">The Sensor Placement Problem</h2>

<p>A concrete instance of OED: you have $N$ candidate sensor locations and a budget of $k \ll N$ sensors. Which $k$ locations should you choose?</p>

<p>Formally, let $C_i \in \mathbb{R}^{1 \times n}$ be the measurement row corresponding to sensor $i$. You seek a subset $\mathcal{S} \subset {1, \ldots, N}$ of size $k$ that minimizes</p>

\[\operatorname{tr}\!\left(\left(\Gamma_\text{prior}^{-1} + \sum_{i \in \mathcal{S}} \sigma_i^{-2} C_i^\top C_i\right)^{-1}\right),\]

<p>where the sensor noises are assumed independent with variances $\sigma_i^2$. Correlated noise requires assembling the covariance for the selected sensor set rather than summing independent contributions.</p>

<p>This is a combinatorial optimization problem: there are $\binom{N}{k}$ possible designs, which for $N = 100$ and $k = 10$ is about $1.7 \times 10^{13}$. Exhaustive search is hopeless.</p>

<p><strong>Greedy methods.</strong> Add sensors one at a time, choosing the sensor that gives the largest decrease in the design criterion. Log-determinant information gain is monotone submodular under common independent-noise assumptions, so the classical greedy algorithm has a $1-1/e$ guarantee for the associated cardinality-constrained maximization problem. A-optimality and E-optimality are not submodular in general, although weaker or problem-specific guarantees are available.</p>

<p><strong>Convex relaxation.</strong> Replace the binary sensor selection indicator by a continuous weight $w_i \in [0, 1]$:</p>

\[\min_{w \in [0,1]^N,\, \mathbf{1}^\top w \leq k} \;\operatorname{tr}\!\left(\left(\Gamma_\text{prior}^{-1} + \sum_{i=1}^N w_i\sigma_i^{-2} C_i^\top C_i\right)^{-1}\right).\]

<p>The relaxed A-optimal objective is convex and admits a semidefinite epigraph formulation. The resulting fractional weights must be rounded or sampled to obtain a binary design, and that step can degrade the relaxed optimum.</p>

<p><strong>Randomized methods.</strong> Sample sensor subsets with probability proportional to their approximate contribution to the information gain. These are scalable and often produce near-optimal designs for large $N$, though without the approximation guarantees of greedy algorithms.</p>

<h2 id="adaptive-and-sequential-design">Adaptive and Sequential Design</h2>

<p>Static design, choosing all measurements before any data is collected, discards information. Sequential design interleaves measurement with inference: after each batch of measurements, the posterior is updated, and the next measurement is chosen based on the current posterior.</p>

<p>The optimal sequential design is the solution of a Bellman equation: at each step, choose the measurement that maximizes the expected information gain over all future steps. This is, in general, computationally intractable for non-trivial numbers of steps.</p>

<p>Approximate strategies that work well in practice include:</p>

<p><strong>Myopic (one-step-ahead) design.</strong> At each step, choose the single measurement that maximally reduces the current criterion. This ignores future steps but is cheap and often effective.</p>

<p><strong>Expected information gain.</strong> For nonlinear models or non-Gaussian priors, closed forms are rarely available. Nested Monte Carlo and variational approximations estimate the expected Kullback-Leibler divergence from posterior to prior. Their computational cost and bias must be assessed for the problem at hand.</p>

<h2 id="application-to-eit">Application to EIT</h2>

<p>In EIT, the measurement protocol specifies which electrode pairs inject current and which pairs measure the resulting voltages. A standard protocol (e.g., adjacent stimulation) applies current between neighbouring electrodes and measures all remaining pairs. But this is a convention, not an optimum.</p>

<p>OED for EIT asks: for a given number of stimulation patterns, which patterns carry the most information about the conductivity? The answer depends on the prior: if you expect a single circular inclusion, the optimal patterns differ from those optimal for a diffuse anomaly.</p>

<p>The preferred protocol depends on the conductivity model, prior, electrode geometry, noise covariance, and design criterion. Adjacent stimulation is therefore not a universal optimum. A quantitative comparison should report all of these choices; percentage improvements from one study do not transfer automatically to another EIT system.</p>

<p>The challenge in applying OED to EIT is that the forward model is nonlinear (the sensitivity of the measurements to the conductivity depends on the conductivity itself), so the information matrix changes as the estimate of the conductivity changes. Adaptive protocols that update the stimulation pattern during the measurement sequence are theoretically superior but require fast online computation. This is an active area of research.</p>

<h2 id="what-i-find-compelling">What I Find Compelling</h2>

<p>OED is fundamentally a question about the structure of information: which features of the unknown can be recovered from which types of data? Understanding this is not only practically useful; it is theoretically clarifying.</p>

<p>In my own work, thinking about optimal designs has changed how I think about the ill-posedness of EIT. The problem is not merely that the forward operator maps many conductivities to similar data; it is that the standard measurement protocols do not extract the information that would distinguish them. This is a design choice, and it can be improved.</p>

<p>The Bayesian framework makes the OED problem tractable and connects it naturally to the regularization and uncertainty quantification methods I described in earlier posts. The design criterion, the prior, the likelihood, and the posterior are all parts of the same coherent probabilistic framework. This integration is one of the things I find most appealing about the Bayesian approach: not just as a method for solving inverse problems, but as a language for thinking about the whole measurement-inference pipeline.</p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Mathematics" /><category term="Inverse Problems" /><category term="Optimal Experimental Design" /><summary type="html"><![CDATA[An introduction to optimal experimental design: how to choose which measurements to take in order to maximally reduce uncertainty about an unknown quantity.]]></summary></entry><entry><title type="html">Optimal Transport: Moving Mass as Efficiently as Possible</title><link href="https://abdgafartunde.github.io/blog/2026/06/22/optimal-transport/" rel="alternate" type="text/html" title="Optimal Transport: Moving Mass as Efficiently as Possible" /><published>2026-06-22T00:00:00+08:00</published><updated>2026-06-22T00:00:00+08:00</updated><id>https://abdgafartunde.github.io/blog/2026/06/22/optimal-transport</id><content type="html" xml:base="https://abdgafartunde.github.io/blog/2026/06/22/optimal-transport/"><![CDATA[<p>Imagine you have a pile of earth and a hole of the same volume. You want to fill the hole by moving the earth, and you want to do it as cheaply as possible, where cost is measured by mass times distance travelled. What is the optimal transport plan?</p>

<p>This is Monge’s problem, posed in 1781. It took nearly 200 years to resolve satisfactorily. The resolution, due primarily to Kantorovich in the 1940s and later refined by Brenier, Benamou, Villani, and many others, produced a theory of extraordinary richness. Optimal transport is now central to analysis, probability, PDEs, image processing, statistics, and machine learning. If you have encountered Wasserstein distances in generative models or geometric measure theory, you have seen one corner of this theory.</p>

<p>I came to optimal transport through two routes: as a tool for comparing distributions in imaging problems, and as a framework for understanding certain data misfit functionals that behave better than the standard $L^2$ norm when the data contains outliers or geometric features. This post is an introduction to the core ideas.</p>

<h2 id="the-monge-problem">The Monge Problem</h2>

<p>Let $\mu$ and $\nu$ be two probability measures on a metric space $\mathcal{X}$, representing respectively the source and target distributions. Think of $\mu$ as the initial pile of earth and $\nu$ as the desired final configuration.</p>

<p>A <strong>transport map</strong> $T : \mathcal{X} \to \mathcal{X}$ pushes $\mu$ forward to $\nu$, meaning that for every measurable set $B$,</p>

\[\nu(B) = \mu(T^{-1}(B)).\]

<p>This is written $T_\sharp \mu = \nu$: the pushforward of $\mu$ under $T$ is $\nu$.</p>

<p>Given a cost function $c(x, y)$ measuring the cost of moving a unit of mass from $x$ to $y$, Monge’s problem is:</p>

\[\inf_{T : T_\sharp \mu = \nu} \int_{\mathcal{X}} c(x, T(x))\, d\mu(x).\]

<p>This is elegant but problematic. The constraint $T_\sharp \mu = \nu$ is nonlinear (it is a PDE constraint on $T$). Worse, a transport map may not exist at all when $\mu$ is a discrete measure: you cannot split a point mass and send it to two different locations using a single map.</p>

<h2 id="the-kantorovich-relaxation">The Kantorovich Relaxation</h2>

<p>Kantorovich’s key insight was to relax the problem. Instead of requiring that every particle at $x$ go to a single destination $T(x)$, allow mass splitting: a particle can be distributed across multiple destinations.</p>

<p>A <strong>transport plan</strong> is a joint probability measure $\pi$ on $\mathcal{X} \times \mathcal{X}$ with marginals $\mu$ and $\nu$:</p>

\[\pi(A \times \mathcal{X}) = \mu(A), \quad \pi(\mathcal{X} \times B) = \nu(B)\]

<p>for all measurable $A, B$. The value $\pi(dx, dy)$ represents the mass transported from $x$ to $y$.</p>

<p>The <strong>Kantorovich problem</strong> is:</p>

\[\inf_{\pi \in \Pi(\mu, \nu)} \int_{\mathcal{X} \times \mathcal{X}} c(x, y)\, d\pi(x, y),\]

<p>where $\Pi(\mu, \nu)$ is the set of all couplings with marginals $\mu$ and $\nu$. This is a <em>linear</em> program over an infinite-dimensional space. It always has a solution (under mild conditions), and its value is called the <strong>optimal transport cost</strong>.</p>

<p>For costs satisfying suitable structural conditions, absolute continuity of $\mu$ can imply that an optimal Kantorovich plan is induced by a map. For the quadratic cost on $\mathbb{R}^d$, Brenier’s theorem gives this conclusion: the optimal plan is supported on the graph of an optimal map $T$, and $\pi = (\operatorname{id},T)_\sharp\mu$.</p>

<h2 id="the-wasserstein-distance">The Wasserstein Distance</h2>

<p>For the cost function $c(x, y) = \lVert x - y \rVert^p$ with $p \geq 1$ on $\mathbb{R}^d$, the optimal transport cost defines the <strong>$p$-Wasserstein distance</strong>:</p>

\[W_p(\mu, \nu) = \left(\inf_{\pi \in \Pi(\mu, \nu)} \int \lVert x - y \rVert^p\, d\pi(x, y)\right)^{1/p}.\]

<p>The resulting $W_p$ is a genuine metric on the space of probability measures with finite $p$-th moment. It metrizes weak convergence of measures together with convergence of moments, which makes it stronger than the weak topology but weaker than total variation.</p>

<p>The case $p = 2$ is especially tractable. By Brenier’s theorem (1991), when $\mu$ is absolutely continuous, the optimal transport map for the quadratic cost is the gradient of a convex function: $T = \nabla \varphi$ for some convex $\varphi : \mathbb{R}^d \to \mathbb{R}$. This characterization connects optimal transport to convex analysis and to the Monge-Ampère PDE:</p>

\[\det(D^2 \varphi(x)) = \frac{\rho_\mu(x)}{\rho_\nu(\nabla \varphi(x))},\]

<p>where $\rho_\mu$ and $\rho_\nu$ are the source and target densities. This formula requires sufficient regularity and positivity of the densities, together with the convexity of $\varphi$.</p>

<p>For $p = 1$, the Wasserstein distance has a dual representation (the Kantorovich-Rubinstein formula):</p>

\[W_1(\mu, \nu) = \sup_{\lVert f \rVert_{\text{Lip}} \leq 1} \int f\, d\mu - \int f\, d\nu,\]

<p>where the supremum is over 1-Lipschitz functions. This formula connects $W_1$ to functional analysis and makes it computable via linear programming duality. It is the form used in Wasserstein GANs, where the discriminator is trained to be a 1-Lipschitz function.</p>

<h2 id="why-w_p-is-better-than-l2-in-imaging">Why $W_p$ is Better Than $L^2$ in Imaging</h2>

<p>The squared $L^2$ distance between two images $f$ and $g$ is</p>

\[\lVert f - g \rVert_{L^2}^2 = \int |f(x) - g(x)|^2\, dx.\]

<p>This measures pointwise amplitude discrepancy. For a translated profile $g(x)=f(x-\varepsilon)$, the $L^2$ distance depends on the overlap between the two profiles and tends to zero as $\varepsilon\to0$ for every $f\in L^2$. It can nevertheless remain large relative to the displacement when narrow, weakly overlapping features are shifted.</p>

<p>The Wasserstein distance records the displacement directly. If $\nu$ is the translation of a probability measure $\mu$ by a vector $a$, then $W_p(\mu,\nu)=\lVert a\rVert$. This can make Wasserstein distances useful for comparing normalized nonnegative signals whose geometry matters.</p>

<p>In seismic full-waveform inversion, suitably constructed optimal-transport misfits can have a wider basin of attraction than an $L^2$ misfit for translated signals. This can reduce cycle skipping in model problems. Seismic traces are signed and oscillatory, so the data must be transformed or compared with an unbalanced or signed transport formulation before a Wasserstein distance is applied.</p>

<h2 id="computing-optimal-transport">Computing Optimal Transport</h2>

<p>For two discrete measures with $n$ support points each, the Kantorovich problem is a linear program with $n^2$ transport variables and $2n$ marginal constraints. The Hungarian algorithm applies to the special assignment problem with uniform atomic masses. General discrete transport uses network-flow, linear-programming, or regularized methods, and storing a dense $n\times n$ cost matrix is already prohibitive for large $n$.</p>

<p><strong>Entropic regularization.</strong> Cuturi’s 2013 paper introduced Sinkhorn’s algorithm for optimal transport by adding an entropy term:</p>

\[\inf_{\pi \in \Pi(\mu, \nu)} \int c\,d\pi
+ \varepsilon\,\operatorname{KL}(\pi\,\|\,\mu\otimes\nu).\]

<p>For strictly positive Gibbs kernels in the discrete setting, the regularized problem has a unique coupling and Sinkhorn iterations compute its scaling factors. Each dense iteration costs $O(n^2)$. The iteration count and regularization bias depend on the cost scale, the desired accuracy, and $\varepsilon$; small values of $\varepsilon$ can cause slow convergence and numerical underflow. Stabilized and multiscale implementations extend the practical range of the method.</p>

<p><strong>Sliced Wasserstein distance.</strong> For measures on $\mathbb{R}^d$ with large $d$, the Wasserstein distance is expensive even with Sinkhorn. The sliced Wasserstein distance averages one-dimensional Wasserstein distances over random projections:</p>

\[\operatorname{SW}_p(\mu,\nu)
= \left(\int_{S^{d-1}} W_p^p((\Pi_\theta)_\sharp\mu,(\Pi_\theta)_\sharp\nu)\,d\theta\right)^{1/p},\]

<p>where $\Pi_\theta$ is projection onto direction $\theta$ and the sphere measure is normalized. For empirical measures with equal weights, each one-dimensional Wasserstein distance is computed by sorting. Sliced Wasserstein is itself a metric under the usual moment assumptions. A Monte Carlo average over finitely many directions approximates the sphere integral.</p>

<h2 id="connections-to-other-areas-i-work-on">Connections to Other Areas I Work On</h2>

<p>Optimal transport appears in several places adjacent to my research.</p>

<p><strong>Data misfit in inverse problems.</strong> When the measured data has structure that the $L^2$ norm poorly captures (signal shifts, geometric features), a Wasserstein-based misfit is an alternative. The gradient of the Wasserstein misfit can be computed via the adjoint method I described in an earlier post, though the non-smoothness of $W_1$ requires care.</p>

<p><strong>Generative priors.</strong> Learned priors for inverse problems often model the distribution of plausible solutions. Wasserstein distances provide a natural way to measure discrepancy between the model’s output distribution and the training distribution, and to formulate the training objective in a way that is less sensitive to mode collapse than KL-divergence-based methods.</p>

<p><strong>Interpolation between measures.</strong> The Wasserstein geodesic (the optimal interpolation between $\mu$ and $\nu$ under the $W_2$ metric) is the displacement interpolation: at time $t \in [0,1]$, the interpolated measure is $(T_t)_{\sharp} \mu$ where $T_t(x) = (1-t)x + tT(x)$. This provides a geometrically meaningful way to interpolate between conductivity distributions in EIT or between imaging data from different patients.</p>

<p>The theory is deep enough to occupy a career, and I am only beginning to understand it. But the connections to problems I care about are direct enough that I expect optimal transport to become a more central part of my toolkit over the next few years.</p>]]></content><author><name>Abd&apos;gafar Tunde Tiamiyu</name></author><category term="Mathematics" /><category term="Optimal Transport" /><category term="Inverse Problems" /><summary type="html"><![CDATA[An introduction to optimal transport, the Wasserstein distance, and why moving probability mass has become a central tool in mathematics, imaging, and machine learning.]]></summary></entry></feed>