<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Emory REU Computational Mathematics for Data Science</title><link>http://www.math.emory.edu/site/cmds-reuret/</link><atom:link href="http://www.math.emory.edu/site/cmds-reuret/index.xml" rel="self" type="application/rss+xml"/><description>Emory REU Computational Mathematics for Data Science</description><generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Fri, 01 Aug 2031 00:00:00 +0000</lastBuildDate><image><url>http://www.math.emory.edu/site/cmds-reuret/media/icon_hu_95a4f94e52276402.png</url><title>Emory REU Computational Mathematics for Data Science</title><link>http://www.math.emory.edu/site/cmds-reuret/</link></image><item><title>Synthetically Rebalancing Healthcare Datasets via Conditional DDPM</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2023-ai-for-healthcare/</link><pubDate>Fri, 01 Dec 2023 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2023-ai-for-healthcare/</guid><description>&lt;!-- https://docs.google.com/document/d/1nmHuz-sI2og6l4DuuRhmWrhOHthWfbb38AnimgdriEo/edit#heading=h.31zpzrlki3gd
-->
&lt;p>This blog post was written by Keira Behal, Jiayi Chen, Caleb Fikes, and Sophia Xiao and published with minor edits. Our work was guided by &lt;a href="../author/yuanzhe-xi/">Dr. Yuanzhe Xi&lt;/a>.In addition to this post, the team has also given a &lt;a href="content/2023_REU_Fair_Health_Presentation.pdf">midterm presentation&lt;/a>, created a &lt;a href="https://youtu.be/eyijGEz9CZg" target="_blank" rel="noopener">poster blitz video&lt;/a>, created a &lt;a href="content/2023_REU_Fair_Health_Poster.pdf">poster&lt;/a>, created some code and wrote a &lt;a href="https://arxiv.org/abs/2310.18430" target="_blank" rel="noopener">paper&lt;/a>.&lt;/p>
&lt;h2 id="background">Background&lt;/h2>
&lt;p>In recent years, machine learning algorithms have become increasingly important in healthcare for tasks like disease prediction, diagnostics, and treatment optimization. However, these algorithms can perpetuate harmful societal biases if the training data contains inherent biases or underrepresentation of certain groups, particularly minority communities. This bias issue is a significant concern in healthcare, where fair and equitable treatment is crucial. Even valuable data sources like Electronic Health Records (EHR), which provide a comprehensive overview of a patient&amp;rsquo;s health history, including diagnoses, treatments, and demographics, can suffer from underrepresentation of certain racial or ethnic minorities. This imbalance can lead to inequitable health outcomes, with minority groups receiving less accurate diagnoses or treatment recommendations due to their underrepresentation in the training data.&lt;/p>
&lt;h2 id="our-approach">Our Approach&lt;/h2>
&lt;p>To address this challenge, we propose &lt;a href="https://arxiv.org/abs/2310.18430" target="_blank" rel="noopener">Minority Class Rebalancing through Augmented Data Generation (McRAGE)&lt;/a>, a novel approach to augment imbalanced medical datasets using samples generated by a deep generative model. The McRAGE process involves training a Conditional Denoising Diffusion Probabilistic Model (CDDPM) capable of generating high-quality synthetic EHR samples from underrepresented classes. We use this synthetic data to augment the imbalanced dataset, achieving a more balanced distribution across all classes. The dataset can be used to train an unbiased machine learning model. Our work intends to promote fair and accurate healthcare predictions, enhancing patient care and supporting equity through data science and artificial intelligence.&lt;/p>
&lt;p>In our project, we focused on data-driven methods of synthetic sample generation through deep generative models because deep generative models aim to capture the underlying data distribution and generate samples that closely resemble the real data so that they can produce high-fidelity synthetic samples that retain the characteristics and variability of the original data. We chose a CDDPM to generate synthetic samples to augment our dataset for the following three reasons:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Stability of Training&lt;/strong>: DDPMs are easier to train than GANs (Generative Adversarial Networks), which are more susceptible to mode collapse and vanishing gradients.&lt;/li>
&lt;li>&lt;strong>High Fidelity Generation&lt;/strong>: diffusion models can produce higher quality samples than other generative models. Many commercial implementations of diffusion models such as Stability AI’s Stable Diffusion and Open AI’s DALL-E 2 have garnered significant funding and public interest.&lt;/li>
&lt;li>&lt;strong>Class-specific Generation&lt;/strong>: The CDDPM offers control over the generation of specific classes of data, whereas conventional DDPM is limited to generating samples that reflect the class distribution of its training set. By utilizing CDDPM, we can use the knowledge gained from the majority groups to improve the quality of generated data for the minority groups.&lt;/li>
&lt;/ol>
&lt;p>The MCRAGE process is both intuitive and theoretically justifiable. The algorithm results in a synthetically rebalanced training set where each “intersectional group,” or a unique combination of sensitive demographic factors, is equally represented. By generating an artificial stratified random sample, the process promotes statistical parity, meaning that classifiers trained on this data will have the same distribution of decisions for each group. This ensures that the classifier&amp;rsquo;s performance is nearly equivalent for each subgroup despite class imbalances in the training data.
The MCRAGE process is proposed under the assumption that such a result holds. In that case, the process results in the nearest approximation to a stratified sample given all the information in the training set. This project aims to empirically test the performance of MCRAGE compared to previous state-of-the-art methods for mitigating dataset imbalance.&lt;/p>
&lt;h2 id="our-contribution">Our Contribution&lt;/h2>
&lt;p>Our approach to conditional generation in the diffusion model is called classifier-free guidance. It involves modifying the denoising update process by incorporating class information. Unlike the classifier guidance method, which relies on a separate classifier for conditional sample generation, our approach integrates the class information directly into the model. In our implementation, we added an extra class embedding to the time embedding in the conventional DDPM. This enables us to incorporate and utilize class information effectively.&lt;/p>
&lt;h2 id="our-results">Our Results&lt;/h2>
&lt;p>While developing our algorithm, we used a subset of MNIST to track our progress. Finally, we applied the algorithm to a tabular EHR dataset called &lt;a href="https://www.kaggle.com/datasets/manishkc06/patient-treatment-classification" target="_blank" rel="noopener">Patient Treatment Classification&lt;/a> with a random forest classifier. The dataset consists of Electronic Health Records collected from a private Hospital in Indonesia. The dataset contains eight laboratory test results for 3309 patients, used to determine patient treatment classification (inpatient care or outpatient care). In both cases, we compared the synthetically balanced datasets of our approach to the balanced dataset and those obtained from the &lt;a href="https://arxiv.org/abs/1106.1813" target="_blank" rel="noopener">SMOTE algorithm&lt;/a>. We also compared the different approaches quantitatively using different fitness metrics and found that the classifier trained on data augmented by the DDPM performed the best.&lt;/p>
&lt;p>These promising results motivate us to conduct the same procedure on a more extensive EHR data set with more features with multiple classes, possibly incorporating a multinomial diffusion.&lt;/p>
&lt;h2 id="interested-in-learning-more">Interested in Learning More?&lt;/h2>
&lt;p>Please see our &lt;a href="content/2023_REU_Fair_Health_Poster.pdf">poster&lt;/a> and &lt;a href="https://arxiv.org/abs/2310.18430" target="_blank" rel="noopener">paper&lt;/a> for more details.&lt;/p></description></item><item><title>A Tensor SVD-based Classification Algorithm Applied to fMRI Data</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2021-tensor/</link><pubDate>Tue, 14 Dec 2021 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2021-tensor/</guid><description>&lt;p>This post was written by Katy Keegan, Yihua Xu, Tanvi Vishwanath, and Vida Jon and published with minor edits. The team was advised by Dr. Elizabeth Newman.
In addition to this post, the team has also given a &lt;a href="https://github.com/EmoryMLIP/emory-reu-ret-website/blob/main/content/projects/2021-tensor/img/_Emory_REU_RET__Summer_2021__Tensor_fMRI_Presentation.pdf" target="_blank" rel="noopener">midterm presentation&lt;/a> , created a &lt;a href="https://github.com/EmoryMLIP/emory-reu-ret-website/blob/main/content/projects/2021-tensor/img/Tensor_fMRI_Poster.pdf" target="_blank" rel="noopener">poster&lt;/a> , published, &lt;a href="https://github.com/elizabethnewman/tensor-fmri" target="_blank" rel="noopener">code&lt;/a>, and written a &lt;a href="https://arxiv.org/abs/2111.00587" target="_blank" rel="noopener">paper&lt;/a>.&lt;/p>
&lt;h2 id="overview-can-we-look-at-a-brain-scan-and-know-what-the-brains-owner-is-thinking">Overview: Can we look at a brain scan and know what the brain&amp;rsquo;s owner is thinking?&lt;/h2>
&lt;p>Our research attempts to use computers to correctly classify brain scans into 2 groups, depending on what they are doing during the scan.&lt;/p>
&lt;h2 id="functional-mri">Functional MRI&lt;/h2>
&lt;p>We use &lt;a href="https://en.wikipedia.org/wiki/Functional_magnetic_resonance_imaging" target="_blank" rel="noopener">brain scans called functional MRIs&lt;/a> that show us which parts of the brain are using more oxygen and are therefore most active. &lt;a href="http://www.cs.cmu.edu/afs/cs.cmu.edu/project/theo-81/www/" target="_blank" rel="noopener">We have fMRIs of test subjects&lt;/a> who, as they are being scanned, are also shown either a picture or a sentence. If computers can classify these study subjects into one of these two categories by only studying their scans, then in a sense we can read their minds.&lt;/p>
&lt;p>fMRIs consist of three-dimensional pixels called voxels, and the data are numbers representing colors. Unlike static MRIs which take a scan at one point in time, fMRIs are repeated every few seconds creating a series of images for each trial. Here is an example of an fMRI of one brain during one trial. The different images are of different slices of the brain. The abbreviations on the right refer to brain regions, for example &amp;ldquo;SMA&amp;rdquo; stands for &amp;ldquo;Supplementary Motor Area&amp;rdquo; located at the top center of the head.&lt;/p>
&lt;img src="img/brain1.jpg" alt="brain1" width="400"/>
&lt;p>If we have three-dimensional fMRI brain voxel data for many patients, multiple scans in sequence, then we need to analyze a quantity of data unwieldy even for modern computers.&lt;/p>
&lt;h2 id="image-classification">Image Classification&lt;/h2>
&lt;p>Image classification is using a computer to figure what what an image represents. Computers can&amp;rsquo;t see images, so they use features of the images that it can understand. For example, we can train a computer to match an image to a numerical digit. Computers learn by training on many images, for example &lt;a href="http://yann.lecun.com/exdb/mnist/" target="_blank" rel="noopener">using the MNIST database of handwritten images&lt;/a>. MNIST contains a wide variety of images that can represent a 0 or a 1:&lt;/p>
&lt;img src="https://user-images.githubusercontent.com/50922545/126396168-5835463f-db60-417b-b4ab-5cc4d6e3b2ef.jpg" width="400" class="aligncenter"/>
&lt;img src="https://user-images.githubusercontent.com/50922545/126396380-f4d0bedb-8a49-455c-bc15-1b73b01a77e5.jpg" width="400"/>
&lt;p>We can construct a &amp;ldquo;basis&amp;rdquo;, which can be thought of as a collection of the most relevant features shared by all of the images belonging to that class.&lt;/p>
&lt;p>We see that the basis for Class 0 shows more curved features, while the basis for Class 1 contains traces of more straight and vertical features. To choose the basis the image more closely matches, we compute a &amp;ldquo;projection&amp;rdquo;, not unlike what you may have calculated with vectors. The larger the projection, the better the match to a particular basis.&lt;/p>
&lt;img src="https://user-images.githubusercontent.com/50922545/127238156-e5b94e20-2853-405b-8483-13dc115565e9.jpg" width="400"/>
&lt;p>The computer tries to classify this test image as either a zero or a one:
&lt;img src="https://user-images.githubusercontent.com/50922545/127212660-d7520639-a8e6-4800-b78f-80901f6b7142.jpg" width="100"/>&lt;/p>
&lt;p>This image represents the projection of the test image onto the basis for numeral one:
&lt;img src="https://user-images.githubusercontent.com/50922545/127212640-afe7cd85-5495-4f43-bc67-3a9a1fe4f9aa.jpg" width="100"/>&lt;/p>
&lt;p>The image projection onto the basis for zero, shows differences and inconsistencies:
&lt;img src="https://user-images.githubusercontent.com/50922545/127212651-8047b39b-5aa1-45f7-b0d6-c5fa1f0b986c.jpg" width="100"/>&lt;/p>
&lt;p>The projection shows that the &amp;ldquo;distance&amp;rdquo; between our test image and our two classes is smaller for Class 1 than for Class 0, and so our method classifies our test image as a 1.&lt;/p>
&lt;p>By learning from a training set of images a computer can examine the data in a new image and figure out which digit it most resembles. Similarly, we are training our computer to learn how to use the data in an fMRI to classify our study subjects into those who are shown an image and those who are shown a sentence.&lt;/p>
&lt;h3 id="tensors-and-singular-value-decomposition">Tensors and Singular Value Decomposition&lt;/h3>
&lt;p>Typically, large data sets like the fMRI voxels are stored in matrices, which have some &lt;a href="https://youtu.be/LlKAna21fLE" target="_blank" rel="noopener">powerful tools for extracting the most relevant components.&lt;/a>&lt;/p>
&lt;p>When we store fMRI data in a matrix we lose important relationships between the data points. For example a computer does not know that a particular voxel representing a part of the brain at one moment is that same part of the brain a few seconds later.&lt;/p>
&lt;p>In our work, we study how to store our data in a tensor which is like matrix but with more than 2 dimensions. Our tensor of fMRIs have a total of 5 dimensions, shown in the figure below. The green slices consist of voxels of the brain in 3 spatial dimensions: x,y,z (yellow). For each trial multiple images are taken over several seconds (blue), and there are multiple trials (red).&lt;/p>
&lt;p>&amp;lt;img width=&amp;ldquo;600&amp;rdquo; alt=&amp;ldquo;fmri_tensors&amp;rdquo;&lt;/p>
&lt;p>When we store our voxel data in tensors instead of matrices, we retain all the information about the 3 dimensional location in the brain, the sequence of the image in time, and the trial. Now we want to decompose our tensor data using a method analogous to the way &lt;a href="https://www.youtube.com/watch?v=DG7YTlGnCEo" target="_blank" rel="noopener">matrices can be decomposed&lt;/a>, so we can use far less data and still extract an accurate prediction of what the subject is thinking. In matrices this is called Singular Value Decomposition; we call our approach tensor-SVD or tSVD.&lt;/p>
&lt;p>Here is a diagram of matrix decomposition using SVD:&lt;/p>
&lt;img width="500" src="https://user-images.githubusercontent.com/50922545/126017121-bd017e2c-7fa1-4d23-8989-c0b69dbbbdf3.jpg">
&lt;p>This what we imagine a tensor SVD would be:&lt;/p>
&lt;img width="500" src="https://user-images.githubusercontent.com/50922545/126017399-7151b4e8-c292-4d34-a1a6-20a7181d6824.png">
&lt;p>You&amp;rsquo;ll notice that we use matrix multiplication in working with matrix SVDs. We are searching for the equivalent tensor multiplication.&lt;/p>
&lt;h2 id="other-applications">Other Applications&lt;/h2>
&lt;p>Our research is not only useful for fMRIs; many datasets have multiple dimensions. Streaming entertaining companies have data on thousands of viewers and what movies they&amp;rsquo;ve watched. Hospitals track thousands of patients, each of whom has had multiple lab tests and other studies. If our research enables us to classify our fMRI subjects, then we may also be able to predict whether someone will want to watch Terminator, or whether a patient is likely to have cancer.&lt;/p>
&lt;h2 id="references">References&lt;/h2>
&lt;p>&lt;a href="https://www.sciencedirect.com/science/article/pii/S0024379515004358" target="_blank" rel="noopener">Tensor tensor products with invertible linear transforms&lt;/a>&lt;/p>
&lt;p>&lt;a href="https://www.pnas.org/content/118/28/e2015851118.short" target="_blank" rel="noopener">Tensor-tensor algebra for optimal representation and compression of multiway data&lt;/a>&lt;/p>
&lt;p>&lt;a href="http://www.kolda.net/publication/koba09/" target="_blank" rel="noopener">Tensor decompositions and applications&lt;/a>&lt;/p>
&lt;p>&lt;a href="https://arxiv.org/pdf/1706.09693.pdf" target="_blank" rel="noopener">Image classification using local tensor singular value decompositions&lt;/a>&lt;/p></description></item><item><title>Generative AI for Structure Discovery in Cardiovascular Digital Twins</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2026-digital-twins/</link><pubDate>Fri, 16 Jan 2026 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2026-digital-twins/</guid><description>&lt;p>&lt;strong>Mentors:&lt;/strong> Dr. Marco Tezzele and Dr. Jimena Martin Tempestti&lt;/p>
&lt;h2 id="overview">Overview&lt;/h2>
&lt;p>A digital twin (DT) is a virtual representation of a physical object that dynamically evolves with its real-world counterpart through sensed data, and provides value through optimal decision-making. Effective DTs rely on accurate mathematical representations of causal dependencies and temporal evolution, often modeled via Probabilistic Graphical Models (PGMs). Traditional approaches frequently assume a static, predefined graph topology (e.g., first-order Markov chains) or rely on simplified expert heuristics, which may fail to capture the rich, long-range dependencies inherent in physiological systems.&lt;/p>
&lt;p>This project investigates how generative AI can automate the discovery of optimal mathematical structures for DTs. Students will develop a framework where a DT is represented as a dynamic PGM. Starting from a fully connected graph encompassing observations, latent digital states, quantities of interest, and actions, students will employ generative methods to prune and optimize the graph topology.&lt;/p>
&lt;h2 id="research-challenges">Research Challenges&lt;/h2>
&lt;p>The research will specifically focus on two &amp;ldquo;New Mathematics&amp;rdquo; challenges:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Topology Learning:&lt;/strong> Analyzing how the graph structure affects the identifiability and estimation of digital states&lt;/li>
&lt;li>&lt;strong>Temporal Memory:&lt;/strong> Investigating whether higher-order Markov chains generated by AI models yield better predictions than standard memoryless models&lt;/li>
&lt;/ol>
&lt;h2 id="validation">Validation&lt;/h2>
&lt;p>The methods will be validated on simulated cardiovascular datasets, aiming to predict patient outcomes and the optimal treatment plan for cardiovascular diseases.&lt;/p></description></item><item><title>AI-Assisted Exploration in Algebra and Number Theory</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2026-algebra-nt/</link><pubDate>Fri, 16 Jan 2026 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2026-algebra-nt/</guid><description>&lt;p>&lt;strong>Mentor:&lt;/strong> Dr. Deependra Singh&lt;/p>
&lt;h2 id="overview">Overview&lt;/h2>
&lt;p>Inspired by the recent work &lt;em>Mathematical exploration and discovery at scale&lt;/em> by Georgiev, Gomez-Serrano, Tao, Wagner, and Google DeepMind, this project explores how LLM-based agents can be systematically integrated into research in algebra and number theory. The central idea is to combine three complementary modes of mathematical work into a single pipeline:&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Computational exploration:&lt;/strong> Using evolutionary and code-based agents (in the spirit of AlphaEvolve) to perform large-scale numerical and symbolic experiments, with the aim of identifying patterns and generating plausible conjectures&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Informal mathematical reasoning:&lt;/strong> Developing insights from experimentation into coherent, human-readable arguments, in the style of extended reasoning systems such as DeepThink&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Formal verification:&lt;/strong> Translating successful arguments into fully rigorous, machine-checked proofs using proof assistants such as Lean, guided by tools similar to AlphaProof&lt;/p>
&lt;/li>
&lt;/ol>
&lt;h2 id="approach">Approach&lt;/h2>
&lt;p>As a first stage, we will apply this pipeline to a set of classical problems in number theory, including results on sums of squares, quadratic forms, and the $u$-invariant of local and global fields. These benchmarks serve both as a training ground for deploying the methodology and as a means of evaluating its strengths and limitations, since the underlying results are well understood.&lt;/p>
&lt;p>Having calibrated the approach on these known cases, we will then turn to carefully selected open problems in less explored settings. In particular, we aim to investigate whether this framework can yield new insights for arithmetic questions over fields beyond the local and global fields, such as semi-global fields (for example, $\mathbb{C}((t))(x)$ and $\mathbb{Q}_p(x)$), where similar questions are less well-understood and systematic exploration may be especially informative.&lt;/p>
&lt;h2 id="prerequisites">Prerequisites&lt;/h2>
&lt;p>Fields, Galois Theory.&lt;/p></description></item><item><title>Joint Analysis of X-ray Ptychography and X-ray Fluorescence Reconstruction</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2025-x-ray/</link><pubDate>Tue, 16 Dec 2025 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2025-x-ray/</guid><description>&lt;p>This project was completed by &lt;strong>Mingke Tian&lt;/strong> and &lt;strong>Eric Zou&lt;/strong> under the mentorship of &lt;strong>Yuanzhe Xi&lt;/strong> as part of the 2025 Emory REU program.&lt;/p>
&lt;p>
&lt;figure id="figure-figure-1-reconstruction-results-from-joint-ptychography-and-fluorescence-framework-source-adapted-from-d-j-vine-et-al-simultaneous-x-ray-fluorescence-and-ptychographic-microscopy-of-cyclotella-meneghiniana-2012">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Reconstruction results" srcset="
/site/cmds-reuret/projects/2025-x-ray/img/reconstruction_hu_9cad732c4fc40a04.webp 400w,
/site/cmds-reuret/projects/2025-x-ray/img/reconstruction_hu_3c4214b234dff13d.webp 760w,
/site/cmds-reuret/projects/2025-x-ray/img/reconstruction_hu_12ea79f3dc5be721.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-x-ray/img/reconstruction_hu_9cad732c4fc40a04.webp"
width="660"
height="327"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Figure 1: Reconstruction results from joint ptychography and fluorescence framework. Source: Adapted from D. J. Vine et al., Simultaneous X-ray fluorescence and ptychographic microscopy of Cyclotella meneghiniana, 2012.
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;h2 id="ptychographic-reconstruction-part">Ptychographic Reconstruction Part&lt;/h2>
&lt;p>In the traditional way, we use separate methods for ptychographic and fluorescence reconstruction. We consider a ptychography experiment where the observed data $d_i$ is modeled as:&lt;/p>
&lt;p>$$
d_i = \vert\mathcal{F}(\mathbf{P}_i \mathbf{z})\vert^2 + \epsilon_i
$$&lt;/p>
&lt;p>where:&lt;/p>
&lt;ul>
&lt;li>$\mathcal{F}$: 2D discrete Fourier transform operator&lt;/li>
&lt;li>$\mathbf{P}_i$: Probe matrix&lt;/li>
&lt;li>$z_i$: the object itself which can be written as $z = x + y_i$&lt;/li>
&lt;/ul>
&lt;h2 id="reconstruction-problem">Reconstruction Problem&lt;/h2>
&lt;p>The reconstruction loss function is formulated as the following:&lt;/p>
&lt;p>$$
\min_{\mathbf{P},\mathbf{z}} \Phi(\mathbf{P},\mathbf{z}) = \frac{1}{2} \sum_{j=1}^{N} \Vert \vert \mathcal{F}(\mathbf{P}_j \mathbf{z})\vert - \sqrt{d_j} \Vert_2^2
$$&lt;/p>
&lt;h2 id="x-ray-fluorescence-reconstruction">X-ray Fluorescence Reconstruction&lt;/h2>
&lt;p>X-ray fluorescence reconstruct the real part of the image $x$ using a deconvolution method:&lt;/p>
&lt;p>$$
\min_{\mathbf{w},\mathbf{P}} \sum_{e=1}^{N_e} \Vert \vert\mathbf{P}\vert^2 * \mathbf{w}_e - D_e \Vert_2^2
$$&lt;/p>
&lt;p>where $\mathbf{w}$ represents elemental concentration maps. $D_e$ corresponds to the experimental flourescence map of element $e$.&lt;/p>
&lt;h2 id="separate-optimization-framework">Separate Optimization Framework&lt;/h2>
&lt;p>Separate optimization of ptychographic and fluorescence reconstruction may lead to the results:&lt;/p>
&lt;p>$$
\min_{\mathbf{P},\mathbf{z}} \Phi(\mathbf{P},\mathbf{z}) + \min_{\mathbf{w},\mathbf{P}} \sum_{e=1}^{N_e} \Vert \vert\mathbf{P}\vert^2 * \mathbf{w}_e - D_e \Vert_2^2
$$&lt;/p>
&lt;p>where $\alpha$ is a scaling parameter balancing between the two objectives.&lt;/p>
&lt;h2 id="joint-optimization-framework">Joint Optimization Framework&lt;/h2>
&lt;p>We propose a simultaneous optimization approach:&lt;/p>
&lt;p>$$
\min_{\mathbf{w},\mathbf{P},\mathbf{z}} \sum_{e=1}^{N_e} \Vert \vert\mathbf{P}\vert^2 * \mathbf{w}_e - D_e \Vert_2^2 + \alpha \sum_{j=1}^{N} \Vert \vert \mathcal{F}(\mathbf{P}_j (\sum_e \mathbf{w}_e + i \beta))\vert - \sqrt{d_j} \Vert_2^2
$$&lt;/p>
&lt;p>By using joint method, the loss function is consistently less than the original loss function from the separate method.&lt;/p>
&lt;h2 id="absorption-coefficient-connection">Absorption Coefficient Connection&lt;/h2>
&lt;p>The absorption coefficient relates to elemental concentrations via:&lt;/p>
&lt;p>$$
\mathbf{z} = \delta + i \beta, \quad \delta = \sum_e \mathbf{w}_e \mu_e
$$&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>This joint framework aims to enhance both ptychographic and fluorescence data processing, leveraging their complementarity for improved reconstruction.&lt;/p>
&lt;h2 id="reference">Reference&lt;/h2>
&lt;p>[1] Deng, J., Vine, D.J., Chen, S. et al. X-ray ptychographic and fluorescence microscopy of frozen-hydrated cells using continuous scanning. Sci Rep 7, 445 (2017). &lt;a href="https://doi.org/10.1038/s41598-017-00569-y" target="_blank" rel="noopener">https://doi.org/10.1038/s41598-017-00569-y&lt;/a>&lt;/p></description></item><item><title>Hybrid Regularization for Random Feature Models</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/</link><pubDate>Fri, 01 Aug 2025 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/</guid><description>&lt;p>This blog post was written by Tony Maccagnano, Sarah Skaggs, Madeline Straus, and Advay Vyas and published with minor edits. The team was advised by &lt;a href="https://www.linkedin.com/in/chelseadrum/" target="_blank" rel="noopener">Dr. Chelsea Drum&lt;/a>. In addition to this website, the team has given a midterm presentation, filmed a &lt;a href="https://youtu.be/wualhZW0JC8" target="_blank" rel="noopener">poster blitz video&lt;/a>, created a &lt;a href="2025REUHybridPosterFinal.pdf">poster&lt;/a>, published &lt;a href="https://github.com/blue-light-house/IPFreeHybridReg" target="_blank" rel="noopener">code&lt;/a>, and worked on a manuscript.&lt;/p>
&lt;h2 id="motivation-and-background">Motivation and Background&lt;/h2>
&lt;p>Image classification is the process by which computer models learn to identify distinguishing features in images in order to recognize and categorize them effectively. As a result, it has many diverse applications in bioinformatics, 3D reconstruction, information extraction, and signal and image processing. In particular, the medical field has seen significant benefits from image classification, which has enabled earlier and more accurate diagnoses by detecting patterns that often elude the human eye and by helping to reduce the rate of misdiagnosis.&lt;/p>
&lt;p>Two main problems within image classification are feature extraction and image labeling. For example, the image below represents photos of the widely studied MNIST dataset which contains handwritten digits 0-9. This dataset serves as a benchmark for training and evaluating classification models before applying them to real-world applications. A well-trained model should be able to efficiently and accurately classify each image as the correct digit before progressing to more complex and diverse classification problems.&lt;/p>
&lt;figure id="figure-sample-images-from-the-mnist-dataset-of-handwritten-digits">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Sample images from the MNIST dataset of handwritten digits." srcset="
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_hu_c0089962e531919d.webp 400w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_hu_ce0c21c1d6317b7c.webp 760w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_hu_aa93bbadb99f7351.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_hu_c0089962e531919d.webp"
width="655"
height="325"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Sample images from the MNIST dataset of handwritten digits.
&lt;/figcaption>&lt;/figure>
&lt;p>Questions that we considered were:&lt;/p>
&lt;ul>
&lt;li>How exactly do computer models learn to identify the contents of an image?&lt;/li>
&lt;li>What happens when there is too much data for the computer system to handle?&lt;/li>
&lt;li>Can we improve the efficiency of existing classification methods?&lt;/li>
&lt;/ul>
&lt;h2 id="our-problem">Our Problem&lt;/h2>
&lt;p>A common and effective approach in image classification is to represent it as a large-scale linear inverse problem of the form $$B \approx AX$$ Here, $B$ is the observed output data (image labels), $A$ is the input image matrix and $X$ is the unknown coefficient matrix we want to approximate. This means we need to determine the input parameters of a system based on the output data. However, many systems are noisy and small disturbances in the label matrix can lead to large differences in the input data, meaning that our problem may be ill-posed. Additionally, when working with large datasets, solving this problem explicitly becomes computationally challenging. One effective strategy to address this is the use of random feature models.&lt;/p>
&lt;h2 id="random-feature-models">Random Feature Models&lt;/h2>
&lt;p>Random feature models transform image data into a higher-dimensional space using randomly generated features, enabling the use of linear methods for classification or regression. A common challenge with this approach is the double descent phenomenon, where test error dips, then spikes, and then falls again as the number of features increases. The spike often appears when the number of features is close to the number of training images, because the system is at the interpolation threshold and becomes ill-conditioned.&lt;/p>
&lt;p>The figure below illustrates the double descent phenomenon when solving the problem directly without regularization. The test error spikes dramatically at $2^{10}$ features, which is also the number of images in this example.&lt;/p>
&lt;figure id="figure-double-descent-phenomenon-in-unregularized-random-feature-models">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Double descent phenomenon in unregularized random feature models." srcset="
/site/cmds-reuret/projects/2025-hybrid-reg/double_descent_hu_e163044faf4161d0.webp 400w,
/site/cmds-reuret/projects/2025-hybrid-reg/double_descent_hu_fedc161d48891020.webp 760w,
/site/cmds-reuret/projects/2025-hybrid-reg/double_descent_hu_534a74b9596de2b3.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/double_descent_hu_e163044faf4161d0.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Double descent phenomenon in unregularized random feature models.
&lt;/figcaption>&lt;/figure>
&lt;p>Instead, we can regularize the problem to stabilize the solution and mitigate the effects of the double descent phenomenon.&lt;/p>
&lt;p>We can either iteratively regularize the problem, meaning we solve it step by step on an increasing subspace, stopping when error is minimal, or we can do variational regularization, which involves solving a slightly different problem that incorporates the weights (complexity) of the solution. However, both of these approaches have their drawbacks. It can be difficult to know when to stop iterative regularization, while variational can take longer to converge and also requires us to provide a regularization parameter $\lambda$, and a bad one can negatively affect the accuracy of the model.&lt;/p>
&lt;p>To get the benefits of these methods while mitigating their downsides, current research has combined the two into something called hybrid regularization. These methods use an iterative method to build the problem subspaces (e.g. Krylov) and apply variational regularization (e.g. Tikhonov) inside those subspaces.&lt;/p>
&lt;h3 id="hybrid-regularization-methods">Hybrid Regularization Methods&lt;/h3>
&lt;p>Previously, in [3], the hybrid LSQR method (available in MATLAB&amp;rsquo;s IR Tools package [2]) was successfully applied for the training of RFMs. Notably, the method:&lt;/p>
&lt;ul>
&lt;li>Avoided the double descent phenomenon&lt;/li>
&lt;li>Performed competitively against existing methods&lt;/li>
&lt;/ul>
&lt;p>However, this method uses computationally expensive inner products. Our proposal is to instead utilize a hybrid LSLU method, introduced in [1], which improves on hybrid LSQR by using a Hessenberg process that&lt;/p>
&lt;ul>
&lt;li>Does not require inner products&lt;/li>
&lt;li>Requires less storage and work&lt;/li>
&lt;/ul>
&lt;p>Through these improvements, this unique approach is advantageous for both mixed precision and high-performance computing scenarios.&lt;/p>
&lt;h2 id="results-and-conclusions">Results and Conclusions&lt;/h2>
&lt;p>For our experiments, we tested various methods for the training of random feature models with the MNIST dataset. We specifically compared the hybrid LSLU method to hybrid LSQR and other standard regularized and unregularized methods.&lt;/p>
&lt;p>We compared test errors against the number of features for different methods on the MNIST dataset, which consists of $70,000$ images ($60,000$ training images, $10,000$ test images) of digits $0, 1, \dots, 9$, where each image is composed of $28 \times 28$ pixels. For this data we can take $n_i = 28 \times 28 = 784$ input features (a gray-scale level for each pixel), and $n_o = 10$ for the number of unique labels (digits). For the following results, we only utilize $2^{10}$ training images. The unregularized method clearly exhibits double descent, while the two simple regularization methods, weight decay and gradient flow have a clear, smooth convergence. However, weight decay and gradient flow require computation of the SVD. Hybrid LSLU and hybrid LSQR converge in a similar manner without requiring SVD.&lt;/p>
&lt;figure id="figure-comparison-of-test-error-vs-number-of-features-for-different-methods">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Comparison of test error vs. number of features for different methods." srcset="
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_features_all_methods_hu_9edfa6fec40ce048.webp 400w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_features_all_methods_hu_ad73a3448417c041.webp 760w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_features_all_methods_hu_b480acc8a0e8387a.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_features_all_methods_hu_9edfa6fec40ce048.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Comparison of test error vs. number of features for different methods.
&lt;/figcaption>&lt;/figure>
&lt;p>In the following figures, we plot test error against number of features for hybrid LSQR and hybrid LSLU on the MNIST data set for different numbers of iterations. We see here that while the accuracy with hybrid LSQR improves with increased iterations, the accuracy of hybrid LSLU is fairly stagnant across most values of features. We do observe that if too many iterations are utilized for hybrid LSLU, then the double descent phenomenon occurs.&lt;/p>
&lt;figure id="figure-test-error-for-hybrid-lsqr-across-different-iterations">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Test error for hybrid LSQR across different iterations." srcset="
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_LSQR_hu_31d27356d48c6010.webp 400w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_LSQR_hu_5fbba964e797373c.webp 760w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_LSQR_hu_b29250090f496ac2.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_LSQR_hu_31d27356d48c6010.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Test error for hybrid LSQR across different iterations.
&lt;/figcaption>&lt;/figure>
&lt;figure id="figure-test-error-for-hybrid-lslu-across-different-iterations">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Test error for hybrid LSLU across different iterations." srcset="
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_LSLU_hu_886f8fb114721fa.webp 400w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_LSLU_hu_b296cfd37069e046.webp 760w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_LSLU_hu_da629ee5b21549ce.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_LSLU_hu_886f8fb114721fa.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Test error for hybrid LSLU across different iterations.
&lt;/figcaption>&lt;/figure>
&lt;p>The following graphs show that hybrid LSLU converges much faster compared to hybrid LSQR in terms of number of iterations.&lt;/p>
&lt;figure id="figure-error-vs-iterations-for-256-features">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Error vs. iterations for 256 features." srcset="
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_256_hu_11e6fce10b7d47dc.webp 400w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_256_hu_1b7f6d4a6c554a9b.webp 760w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_256_hu_786dfd9000224f59.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_256_hu_11e6fce10b7d47dc.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Error vs. iterations for 256 features.
&lt;/figcaption>&lt;/figure>
&lt;figure id="figure-error-vs-iterations-for-1024-features">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Error vs. iterations for 1024 features." srcset="
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_1024_hu_eaa25b64ba9fb752.webp 400w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_1024_hu_12c3b27709ef3930.webp 760w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_1024_hu_f9f32505bdb99dce.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_1024_hu_eaa25b64ba9fb752.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Error vs. iterations for 1024 features.
&lt;/figcaption>&lt;/figure>
&lt;figure id="figure-error-vs-iterations-for-4096-features">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Error vs. iterations for 4096 features." srcset="
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_4096_hu_8c07fbbd2fc1a01e.webp 400w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_4096_hu_16806d000c70908d.webp 760w,
/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_4096_hu_19dd39ea7db1448a.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2025-hybrid-reg/MNIST_error_iterations_4096_hu_8c07fbbd2fc1a01e.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Error vs. iterations for 4096 features.
&lt;/figcaption>&lt;/figure>
&lt;p>Our experiments show that hybrid LSLU converges much faster compared to hybrid LSQR in terms of number of iterations. Thus, not only is hybrid LSLU cheaper (in terms of computational cost) per iteration than hybrid LSQR, but it also requires less iterations to achieve comparable accuracy. We plan to investigate how to optimally stop hybrid LSLU in order to ensure the avoidance of the double descent phenomenon. In our future research we also hope to apply the randomized sketch-and-solve version of hybrid LSLU, which has the potential to address some of the issues we currently see.&lt;/p>
&lt;h2 id="our-activities">Our Activities&lt;/h2>
&lt;h3 id="weeks-1-and-2">Weeks 1 and 2&lt;/h3>
&lt;p>We did literature review on previous and current research. Using IRTools, we replicated the results of the pre-existing papers on hybrid LSQR by implementing random feature model projections.&lt;/p>
&lt;h3 id="week-3">Week 3&lt;/h3>
&lt;p>We composed our midterm presentation and began experimenting with the hybrid LSLU function and how it compared to hybrid LSQR.&lt;/p>
&lt;h3 id="week-4">Week 4&lt;/h3>
&lt;p>We worked on the midterm presentation and then researched different ways to optimize the regularization parameter λ.&lt;/p>
&lt;h3 id="week-5">Week 5&lt;/h3>
&lt;p>We looked more at double descent and different types of GCV in order to optimize LSLU with RFMs. We also started on the webpage draft.&lt;/p>
&lt;h3 id="week-6">Week 6&lt;/h3>
&lt;p>We put together our poster blitz video and started work on the manuscript while also working on getting results for the poster and manuscript. We researched more theory in order to properly explain our ideas.&lt;/p>
&lt;h3 id="week-7">Week 7&lt;/h3>
&lt;p>We put together our poster for the poster presentation and finalized our results and research. We implemented LSLU on MNIST and CIFAR10 and compared the test error to unregularized, weight decay, gradient flow and LSQR. We started working on the manuscript in order to get a rough draft done quickly.&lt;/p>
&lt;h3 id="week-8">Week 8&lt;/h3>
&lt;p>We presented our poster, finished up the manuscript, and cleaned up our code.&lt;/p>
&lt;h2 id="more-about-the-team">More about the Team&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Tony Maccagnano&lt;/strong> is a rising junior at Michigan Technological University, double majoring in Mathematics and Computer Science. He plans to go to graduate school for algebra or combinatorics.&lt;/li>
&lt;li>&lt;strong>Sarah Skaggs&lt;/strong> is a rising junior at Virginia Tech double majoring in Computational Modeling and Data Analytics and Mathematics. She plans to go to graduate school for applied math and her dream career is to use math and data science to advance medical technology. Outside of academics, she loves to play pickleball, adventure, play board games and spend time with friends and family.&lt;/li>
&lt;li>&lt;strong>Madeline Straus&lt;/strong> is a rising junior at Tufts University double majoring in Applied Mathematics and History, with plans to attend graduate school for math, and is particularly interested in numerical analysis and tomography. In her free time, she likes to read, compete at trivia, and play board games.&lt;/li>
&lt;li>&lt;strong>Advay Vyas&lt;/strong> is a rising sophomore at the University of Texas at Austin double majoring in Statistics &amp;amp; Data Science and Mathematics. He hopes to pursue a PhD in Statistics and then work on new ideas in machine learning and its applications. In his free time, he loves to read novels, bike, and watch sports.&lt;/li>
&lt;/ul>
&lt;h2 id="references-and-further-reading">References and Further Reading&lt;/h2>
&lt;ol>
&lt;li>Ariana N Brown et al. &amp;ldquo;Inner Product Free Krylov Methods for Large-Scale Inverse Problems&amp;rdquo;. In: arXiv preprint arXiv:2409.05239 (2024).&lt;/li>
&lt;li>Silvia Gazzola, Per Christian Hansen, and James G Nagy. &amp;ldquo;IR Tools: a MATLAB package of iterative regularization methods and large-scale test problems&amp;rdquo;. In: Numerical Algorithms 81 (2019), pp. 773–811.&lt;/li>
&lt;li>Kelvin Kan, James Nagy, and Lars Ruthotto. &amp;ldquo;Hybrid Regularization Methods Achieve Near-Optimal Regularization in Random Feature Models&amp;rdquo;. In: Submitted to Transactions on Machine Learning Research (2024).&lt;/li>
&lt;li>Ali Rahimi and Benjamin Recht. &amp;ldquo;Random Features for Large-Scale Kernel Machines&amp;rdquo;. In: Advances in Neural Information Processing Systems. Vol. 20. 2007, pp. 1177–1184&lt;/li>
&lt;/ol></description></item><item><title>Inverse problems and uncertainty quantification for the control of dynamical systems</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2025-uq-control/</link><pubDate>Fri, 01 Aug 2025 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2025-uq-control/</guid><description>&lt;p>&lt;em>By Nina Chafee, Micah Chandler, Benjamin Lebdaoui, and Meredith Orton&lt;/em>&lt;/p>
&lt;p>&lt;em>Advised by Dr. John Darges&lt;/em>&lt;/p>
&lt;h2 id="how-can-we-maximize-our-harvest-without-endangering-the-local-wolf-and-bunny-populations">How can we maximize our harvest without endangering the local wolf and bunny populations?&lt;/h2>
&lt;h3 id="introduction">Introduction&lt;/h3>
&lt;p>We&amp;rsquo;re farmers, and we have a big problem! Every spring, the rabbit population skyrockets, decimating our crops, which means we don&amp;rsquo;t have quite enough food to feed everyone. We could douse the crops in rabbit-killing pesticides, but unfortunately, we actually care about the environment and its inhabitants&amp;mdash;particularly the local population of wolves. Negatively impacting the rabbit population too much depletes their food source, and could drive both populations to extinction. So, as much as possible, how can we preserve the local ecosystem &lt;em>and&lt;/em> reduce the bunny boom? You may find this video illuminating: &lt;a href="https://drive.google.com/file/d/1Trk_DmEuz1xbErP1mdDC6G8aOCIjqAA1/view?usp=sharing" target="_blank" rel="noopener">Poster Blitz Video&lt;/a>&lt;/p>
&lt;h3 id="very-educated-guessing">Very Educated Guessing&lt;/h3>
&lt;p>To predict the impact of possible interventions, we first need to understand the dynamics of the system. We can do this by modeling the population dynamics between the bunnies and the wolves with the Lotka-Volterra Model:&lt;/p>
&lt;p>$$\frac{dx}{dt} = \theta_1 x - \theta_2xy$$&lt;/p>
&lt;p>$$\frac{dy}{dt} = -\theta_3 y + \theta_3 xy$$&lt;/p>
&lt;p>$$x(0) = x_0, \quad y(0) = y_0,$$&lt;/p>
&lt;p>where $\theta_1,\theta_2,\theta_3,\theta_4, x_0, y_0$ are the parameters that govern the dynamics of this system.&lt;/p>
&lt;p>Of course, in the real world, we don&amp;rsquo;t know what these governing parameters are! All we have access to is noisy and often incomplete data. From that data, we need to perform &lt;strong>parameter estimation&lt;/strong> to obtain the parameters that will make the model reproduce the data we&amp;rsquo;ve observed (&lt;em>psst&lt;/em> this is an inverse problem*).&lt;/p>
&lt;h4 id="deterministic-approach">Deterministic Approach&lt;/h4>
&lt;p>We assume that our parameters are fixed unknowns. We want to minimize the distance between the output of our model and the data we are observing, measured with the Ordinary Least Squares (OLS) Estimator:&lt;/p>
&lt;p>$$\theta_{OLS} = \text{argmin}&lt;em>{\theta}\sum&lt;/em>{i= 1}^N(F(x_i;\theta)-d_i)^2,$$&lt;/p>
&lt;p>where $\theta$ is parameter vector $[\theta_1,\theta_2,\theta_3,\theta_4,x_0,y_0]$, $F(x;\theta)$ is our model, and $d$ is our data.&lt;/p>
&lt;p>We use data 20 years of local bunny and wolf populations. Note that we use synthetic data, so we know the true parameters for this experiment.&lt;/p>
&lt;img width="800" height="300" alt="Frequentist Results" src="img/freq.png">
&lt;h4 id="bayesian-approach">Bayesian Approach&lt;/h4>
&lt;p>We&amp;rsquo;ll assume our parameters are random variables, and we want to construct posterior distributions for our parameters; that is, what parameters are likely and how likely are they?&lt;/p>
&lt;p>Using &lt;strong>Bayes&amp;rsquo; theorem&lt;/strong>, we combine prior knowledge with observed data:&lt;/p>
&lt;p>$$p(\theta\mid data) = p(data\mid\theta) \cdot p(\theta)$$&lt;/p>
&lt;p>Where:&lt;/p>
&lt;ul>
&lt;li>$p(\theta\mid data)$ = posterior distribution (what we want)&lt;/li>
&lt;li>$p(data\mid \theta)$ = likelihood (how well parameters explain data)&lt;/li>
&lt;li>$p(\theta)$ = prior distribution (our initial beliefs about parameters)&lt;/li>
&lt;/ul>
&lt;p>This gives us:&lt;/p>
&lt;ul>
&lt;li>Robustness to noise&lt;/li>
&lt;li>Prior knowledge incorporation&lt;/li>
&lt;li>Full uncertainty propagation&lt;/li>
&lt;/ul>
&lt;p>We do this through Markov Chain Monte Carlo (MCMC) Methods. Since we can&amp;rsquo;t directly calculate $p(\theta \mid data)$, MCMC creates a &amp;ldquo;chain&amp;rdquo; of parameter samples that eventually converges to the true posterior distribution.&lt;/p>
&lt;p>&lt;strong>Metropolis-Hastings Algorithm&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Proposal distribution: $q(\theta_j , \theta_{j-1})$&lt;/li>
&lt;li>Target distribution: $\Pi_{post}(\theta \mid d)$&lt;/li>
&lt;/ul>
&lt;p>Steps:&lt;/p>
&lt;ol>
&lt;li>
&lt;p>$\theta_j$ , sample $\theta^* \sim q(\theta_j)$&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Evaluate quality of $\theta^*$ compared to $\theta_j$ using $\Pi_{post}$&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Based on evaluation accept $\theta_{j+1} = \theta^*$ or reject $\theta_{j+1} = \theta_j$&lt;/p>
&lt;/li>
&lt;/ol>
&lt;p>Now we predict the future behavior of the bunny and wolf populations for the next 20 years to model future control strategies, and we can construct confidence and prediction intervals to quantify uncertainty from the noisy data:&lt;/p>
&lt;img width="600" height="300" alt="Bayesian Results" src="img/bayes.png">
&lt;h3 id="taking-back-the-farm">Taking back the farm&lt;/h3>
&lt;p>&lt;strong>Goal:&lt;/strong> limit the bunny boom in the spring without hurting the ecosystem! We can do this by driving the system of rabbits and wolves towards an equilibrium which will smooth population fluxuations without forcing either population to extinction.&lt;/p>
&lt;p>Controls, denoted $\alpha(t)$, are actions that we take to change the system. The &lt;em>cost function&lt;/em> measures how unfavorable controls are. The &lt;strong>optimal control&lt;/strong> is the one that minimizes the cost function:&lt;/p>
&lt;p>$$u(t) = \text{argmin} C_{x,t}(\alpha).$$&lt;/p>
&lt;p>The cost function $C_{x,t}(\alpha)$ is often the sum of two parts: a measure of the control&amp;rsquo;s success $g(x(T))$ at the end state, and an integral measuring the control&amp;rsquo;s running cost.&lt;/p>
&lt;p>$$C_{x,t}(\alpha) = g(x(T)) + \int_{t_0}^{T} L(x(t), \alpha(t), t)$$&lt;/p>
&lt;p>While there are a few approaches to control Lotka-Volterra systems, we chose removal and addition of wolves to the system.&lt;/p>
&lt;p>$$\frac{dx}{dt} = ax - bxy$$&lt;/p>
&lt;p>$$\frac{dy}{dt} = -cy + dxy - u(t).$$&lt;/p>
&lt;p>Cost function:&lt;/p>
&lt;p>$$C(\alpha) = (p(T) - \frac{\theta_3}{\theta_4})^2 + (r(T) - \frac{\theta_1}{\theta_2})^2 + \int_0^T\alpha(t)^2dt$$&lt;/p>
&lt;p>To propagate uncertainty, we run optimal control on each sample generated with MCMC, so we can create a posterior distribution of the optimal controls as well. We can then create confidence intervals with the standard deviation at each time point across some 5,000 samples.&lt;/p>
&lt;p>&lt;strong>Pseudo Spectral Method&lt;/strong>&lt;/p>
&lt;p>We approximate our control and states using orthonormal Legendre polynomials, then substitute these approximations into our cost function and dynamical constraints to turn our ODE system into an algebraic one. Once this is done, we can use Gauss quadrature to optimize the new cost function, subject to the new dynamical constraints.&lt;/p>
&lt;img width="1138" height="283" alt="Pseudospectral Control Results" src="img/psuedo_control.png">
&lt;p>&lt;strong>Pontryagin Maximal Principle&lt;/strong>&lt;/p>
&lt;p>After constructing the Hamiltonian, we formulate the optimal control in terms of the costate variable, $\lambda$. We can then use the Shooting Method which utilizes an initial guess, $\lambda_0$, then uses root-finding to find the right value for $\lambda$ that satisfies the boundary conditions.&lt;/p>
&lt;img width="1138" height="283" alt="PMP Method Results" src="img/pmp_control.png">
&lt;h3 id="so-what-did-we-do">So what did we do?&lt;/h3>
&lt;p>We have created a full pipeline to take in noisy data, estimate the parameters of a model, derive and implement an optimal control for the system, and propagate uncertainty to quantify the confidence in our decision-making.&lt;/p>
&lt;p>Poster: &lt;a href="Poster.pdf">Download the Poster (PDF)&lt;/a>&lt;/p>
&lt;p>Slides: &lt;a href="Midterm_Presentation.pdf">Download the Midterm Presentation (PDF)&lt;/a>&lt;/p></description></item><item><title>Model Aware and Data Driven Inference</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2025-model-aware/</link><pubDate>Tue, 01 Jul 2025 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2025-model-aware/</guid><description>&lt;p>This blog post was written by &lt;a href="https://www.linkedin.com/in/alexanderdelise/" target="_blank" rel="noopener">Alexander DeLise&lt;/a>, &lt;a href="https://www.linkedin.com/in/kyle-loh-a2a3272a9/" target="_blank" rel="noopener">Kyle Loh&lt;/a>, &lt;a href="https://www.linkedin.com/in/krish-patel-1a8804224/" target="_blank" rel="noopener">Krish Patel&lt;/a>, and &lt;a href="https://www.linkedin.com/in/meredithcteague/" target="_blank" rel="noopener">Meredith Teague&lt;/a> and published with minor edits. The team was advised by Andrea Arnold and Matthias Chung. In addition to this post, the team has also created a &lt;a href="MADDI-Poster.pdf">poster&lt;/a> and filmed a &lt;a href="https://youtu.be/sEYCTWnZ3pM?si=L2fmIEY2AUwPl54Z" target="_blank" rel="noopener">poster blitz video&lt;/a>.&lt;/p>
&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>Inverse problems arise naturally whenever we want to recover hidden information from indirect, often noisy measurements. Mathematically, an inverse problem usually takes the form:&lt;/p>
&lt;p>$$
\mathbf{y} = \mathbf{F}(\mathbf{x}) + \varepsilon,
$$&lt;/p>
&lt;p>where:&lt;/p>
&lt;ul>
&lt;li>$\mathbf{x}$ represents unknown parameters we aim to determine,&lt;/li>
&lt;li>$\mathbf{F}$ is a known process (forward operator) that describes how these parameters produce observable data,&lt;/li>
&lt;li>$\mathbf{y}$ are the observed measurements,&lt;/li>
&lt;li>$\varepsilon$ is noise or error in the measurement.&lt;/li>
&lt;/ul>
&lt;h2 id="why-inverse-problems-are-difficult">Why Inverse Problems are Difficult&lt;/h2>
&lt;p>Inverse problems are critically important because their solutions drive decisions across science and engineering. However, solving inverse problems directly is notoriously challenging and often impossible due to several inherent issues:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Ill-posedness:&lt;/strong> Solutions may not exist, be unique, or depend continuously on the input data.&lt;/li>
&lt;li>&lt;strong>Noise Sensitivity:&lt;/strong> Small errors or noise in measurements can lead to significantly inaccurate solutions.&lt;/li>
&lt;li>&lt;strong>Computational Difficulty:&lt;/strong> Direct inversion of the operator $\mathbf{F}$ is frequently computationally infeasible or unstable.&lt;/li>
&lt;/ul>
&lt;p>Because of these difficulties, traditional methods often fail, motivating the search for more robust and insightful approaches.&lt;/p>
&lt;div style="text-align: center;">
&lt;img src="img1.png" alt="Inverse Problem Diagram" width="75%">
&lt;/div>
&lt;br>
&lt;h2 id="neural-networks-as-a-potential-solution-and-the-black-box-issue">Neural Networks as a Potential Solution and the &amp;ldquo;Black Box&amp;rdquo; Issue&lt;/h2>
&lt;p>Machine learning, particularly neural networks, has emerged as a promising solution. Neural networks can implicitly learn complex relationships from data, providing accurate approximations without needing explicit inversion.&lt;/p>
&lt;p>However, a major drawback is their &amp;ldquo;black box&amp;rdquo; nature: they provide little to no insight into why and how they work. In fields where reliability and interpretability are crucial, this lack of understanding severely limits their practical adoption.&lt;/p>
&lt;h2 id="our-approach-linear-encoder-decoder-networks">Our Approach: Linear Encoder-Decoder Networks&lt;/h2>
&lt;p>To address this interpretability gap, we investigate &lt;strong>linear encoder-decoder (ED) networks&lt;/strong>. Linear ED networks simplify neural network architecture while maintaining the capacity to model inverse problems effectively:&lt;/p>
&lt;ul>
&lt;li>$\mathbf{E}$ - &lt;strong>Encoder:&lt;/strong> Compresses the input data into a lower-dimensional representation. This is our latent variable $\mathbf{z}$&lt;/li>
&lt;li>$\mathbf{D}$ - &lt;strong>Decoder:&lt;/strong> Attempts to reconstruct the original data from this reduced representation.&lt;/li>
&lt;/ul>
&lt;p>The linearity of these networks allows us to derive clear mathematical expressions and theoretical insights.&lt;/p>
&lt;div style="text-align: center;">
&lt;img src="img2.png" alt="Encoder-Decoder Network Diagram" width="75%">
&lt;/div>
&lt;br>
&lt;h2 id="how-do-we-decide-whats-best">How Do We Decide What&amp;rsquo;s Best?&lt;/h2>
&lt;p>Choosing a way to measure success is crucial. We use &lt;strong>Bayes risk minimization&lt;/strong> because it explicitly minimizes the expected reconstruction error:&lt;/p>
&lt;p>$$
\min_{\operatorname{rank}(\mathbf{A}) \leq r} \ \mathbb{E} \left|\mathbf{A}Y - X \right|_2^2 ,
$$&lt;/p>
&lt;p>where:&lt;/p>
&lt;ul>
&lt;li>$\mathbf{A}$ is the mapping (network) we want to learn,&lt;/li>
&lt;li>$Y$ are the observed measurements,&lt;/li>
&lt;li>$X$ is the true unknown data.&lt;/li>
&lt;/ul>
&lt;p>Note that $X, Y$ are now &lt;strong>random variables&lt;/strong>, and their realizations can be thought of as our datapoints. Bayes&amp;rsquo; Risk Minimization helps us systematically find solutions that are expected to perform best.&lt;/p>
&lt;h2 id="our-results-in-action">Our Results in Action&lt;/h2>
&lt;p>Let&amp;rsquo;s take a look at two common scenarios in inverse problems and compare how our theoretical optimal mappings perform against the learned mappings by our ED networks.&lt;/p>
&lt;h3 id="linear-denoising">Linear Denoising&lt;/h3>
&lt;h4 id="the-theory">The Theory&lt;/h4>
&lt;p>The &lt;strong>Linear Denoising Problem&lt;/strong> is one of the simplest yet most informative cases of an inverse problem. Here, we aim to recover the original signal $X$ from noisy observations $Y$ that are direct perturbations of $X$, i.e., there is no intermediate transformation like a forward operator. Mathematically, we assume&lt;/p>
&lt;p>$$
Y = X + \mathcal{E},
$$&lt;/p>
&lt;p>where $\mathcal{E}$ is a random noise term.&lt;/p>
&lt;p>Our goal is to find a &lt;strong>low-rank linear map&lt;/strong> $\mathbf{A}$ that minimizes the expected squared reconstruction error between the predicted and true signals:&lt;/p>
&lt;p>$$
\min_{\operatorname{rank}(\mathbf{A}) \leq r} \ \mathbb{E} \ \left| \mathbf{A} Y - X \right|_2^2.
$$&lt;/p>
&lt;p>This optimization is framed through the lens of &lt;strong>Bayes&amp;rsquo; Risk Minimization&lt;/strong>, which allows us to derive closed-form solutions for the optimal map. Just like in more complex inverse problems, we assume that $X$ and $\mathcal{E}$ are random variables with finite second moments, and we use symmetric matrix decompositions to analyze the structure of their variability.&lt;/p>
&lt;div style="text-align: center;">
&lt;img src="img4.png" alt="Noisy Data Example" width="75%">
&lt;/div>
&lt;br>
&lt;p>To do this, we compute the second-moment matrices:&lt;/p>
&lt;ul>
&lt;li>$\mathbf{\Gamma}_X = \mathbb{E}[XX^\top]$ for the clean data,&lt;/li>
&lt;li>$\mathbf{\Gamma}_{Y} = \mathbb{E}[YY^\top]$ for the noise.&lt;/li>
&lt;/ul>
&lt;p>These are decomposed symmetrically, for instance via &lt;a href="https://en.wikipedia.org/wiki/Matrix_decomposition#Cholesky_decomposition" target="_blank" rel="noopener">Cholesky decomposition&lt;/a> or &lt;a href="https://en.wikipedia.org/wiki/Matrix_decomposition#Eigendecomposition" target="_blank" rel="noopener">eigendecomposition&lt;/a>, into:&lt;/p>
&lt;p>$$
\mathbf{\Gamma}_{X} = \mathbf{L}_{X} \mathbf{L}_{X}^\top, \quad \mathbf{\Gamma}_{Y} = \mathbf{L}_{Y} \mathbf{L}_{Y}^\top,
$$&lt;/p>
&lt;p>where $\mathbf{L}_{X}$ and $\mathbf{L}_{Y}$ need not be full rank. The solution to the linear denoising optimization problem is given by&lt;/p>
&lt;p>$$
\mathbf{A}_{\text{opt}}^r = \left( \mathbf{\Gamma}_X \mathbf{L}_Y^{\dagger , \top} \right)_r \mathbf{L}_Y^{\dagger},
$$&lt;/p>
&lt;p>where $(\cdot)_r$ denotes the &lt;a href="https://en.wikipedia.org/wiki/Singular_value_decomposition#Low-rank_matrix_approximation" target="_blank" rel="noopener">rank-$r$ truncated Singular Value Decomposition&lt;/a> (SVD) of a matrix. This provides a clean analytic expression for the best low-rank denoiser under the assumed distributions.&lt;/p>
&lt;h4 id="the-experiment">The Experiment&lt;/h4>
&lt;p>In our experimental setup, we consider biomedical image data drawn from the &lt;a href="https://medmnist.com/" target="_blank" rel="noopener">MedMNIST dataset&lt;/a>. Each image is first vectorized to form a column vector $\mathbf{x}_j \in \mathbb{R}^{784}$. To simulate the denoising problem, we add Gaussian white noise to each vector:&lt;/p>
&lt;ul>
&lt;li>The noise is drawn independently from a zero-mean Gaussian distribution with standard deviation $\sigma = 0.05$.&lt;/li>
&lt;li>This yields the observed measurement:&lt;/li>
&lt;/ul>
&lt;p>$$
\mathbf{y}_j = \mathbf{x}_j + \varepsilon_j, \quad \text{where } \varepsilon_j \sim \mathcal{N}\left( \mathbf{0}, \ 0.05^2 \cdot \mathbf{I} \right).
$$&lt;/p>
&lt;p>We then concatenate these entries into the matrices $\mathbf{X}, \mathbf{Y}$ for our data and observations, respectively. This setup allows us to test how well both learned and theoretical low-rank mappings can remove noise and recover the original image signals.&lt;/p>
&lt;h4 id="the-results">The Results&lt;/h4>
&lt;p>We compare our &lt;strong>Bayes-optimal mappings&lt;/strong> $\mathbf{A}_{\text{opt}}^r$ against &lt;strong>learned linear encoder-decoder mappings&lt;/strong> $\mathbf{A}_{\text{learn}}^r$ trained using gradient descent to minimize empirical reconstruction error. As expected, the theoretical mappings consistently outperform the learned ones, particularly at &lt;strong>low ranks&lt;/strong>, where the model must compress the data most aggressively.&lt;/p>
&lt;div style="text-align: center;">
&lt;img src="img5.gif" alt="Linear Denoising Results Animation" width="75%">
&lt;/div>
&lt;br>
&lt;p>This demonstrates that even in a simplified setting, our analytic solutions are highly efficient at recovering the true signal from noisy observations using only a small number of latent features.&lt;/p>
&lt;h3 id="inverse-end-to-end-problem">Inverse End-to-End Problem&lt;/h3>
&lt;h4 id="the-theory-1">The Theory&lt;/h4>
&lt;p>The &lt;strong>Inverse End-to-End Problem&lt;/strong> is a more general formulation of the data denoising problem, and it refers to the task of recovering the original unknown parameters $X$ from indirect, noisy observations $Y$ that have been passed through a known forward process $\mathbf{F}$ and perturbed by some noise $\mathcal{E}$. Mathematically, we observe&lt;/p>
&lt;p>$$
Y = \mathbf{F} X + \mathcal{E},
$$&lt;/p>
&lt;p>and seek to find a low-rank linear map $\mathbf{A}$ that best approximates the inverse mapping, i.e., recovers $X$ from $Y$ by minimizing the expected reconstruction error:&lt;/p>
&lt;p>$$
\min_{\operatorname{rank}(\mathbf{A}) \leq r} \ \mathbb{E} \ \left| \mathbf{A} Y - X \right|_2^2.
$$&lt;/p>
&lt;div style="text-align: center;">
&lt;img src="img6.png" alt="Inverse End-to-End Problem Diagram" width="75%">
&lt;/div>
&lt;br>
&lt;p>Our theoretically optimal mapping has the form&lt;/p>
&lt;p>$$
\mathbf{A}_{\text{opt}}^r = \left( \mathbf{\Gamma}_X \mathbf{F}^\top \mathbf{L}_Y^{\dagger , \top} \right)_r \mathbf{L}_Y^{\dagger},
$$&lt;/p>
&lt;p>where $\mathbf{\Gamma}_X$ is the second-moment matrix of the random variable $X$, $\mathbf{L}_Y$ comes from a symmetric decomposition of the second-moment matrix of the random variable $Y$ (i.e. $\mathbf{\Gamma}_Y = \mathbf{L}_Y \mathbf{L}_Y^\top$), and $(\cdot)_r$ denotes the rank-$r$ truncated SVD of a matrix, as mentioned before.&lt;/p>
&lt;h4 id="the-experiment-1">The Experiment&lt;/h4>
&lt;p>In this numerical experiment, we define the forward process $\mathbf{F}$ as a full rank &lt;strong>Gaussian blur operator&lt;/strong>. Specifically:&lt;/p>
&lt;ul>
&lt;li>The blur is implemented using a $5 \times 5$ &lt;strong>Gaussian kernel&lt;/strong> with standard deviation $\sigma = 1.0$.&lt;/li>
&lt;li>This is followed by the addition of &lt;strong>Gaussian white noise&lt;/strong> with standard deviation $\sigma = 0.05$.&lt;/li>
&lt;/ul>
&lt;p>Again, our data comes from the MedMNIST dataset. Our observed data is thus generated by:&lt;/p>
&lt;p>$$
\mathbf{y}_j = \mathbf{F}\mathbf{x}_j + \varepsilon_j, \quad \text{where } \varepsilon_j \sim \mathcal{N}\left( \mathbf{0}, \ 0.05^2 \cdot \mathbf{I} \right).
$$&lt;/p>
&lt;p>Again, we then concatenate these entries into the matrices $\mathbf{X}, \mathbf{Y}$ for our data and observations, respectively.&lt;/p>
&lt;h4 id="the-results-1">The Results&lt;/h4>
&lt;p>The animation below provides a clear demonstration of how the optimal rank-$r$ mapping $\mathbf{A}_{\text{opt}}^r$ consistently outperforms the learned $\mathbf{A}_{\text{learn}}^r$, especially at very low ranks.&lt;/p>
&lt;div style="text-align: center;">
&lt;img src="img6.gif" alt="Inverse End-to-End Results Animation" width="75%">
&lt;/div>
&lt;br>
&lt;p>Once again, this means that, even at very low ranks, our Bayes&amp;rsquo; Risk-derived mappings recover the original parameters $X$ more accurately than the learned mappings, and hence do a better job of extracting what we truly care about from noisy, indirect measurements.&lt;/p>
&lt;h2 id="acknowledgements">Acknowledgements&lt;/h2>
&lt;p>This work was conducted as part of the NSF REU Computational Mathematics for Data Science program at Emory University. The authors acknowledge NSF DMS-2349534 for support of this research.&lt;/p></description></item><item><title>Automatic Differentiation for Image Registration</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2024-registration/</link><pubDate>Mon, 08 Jul 2024 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2024-registration/</guid><description>&lt;p>This blog post was written by Cash Cherry, Warin Watson, and Rachelle Lang and published with minor edits. The team was advised by &lt;a href="https://www.math.emory.edu/~lruthot/" target="_blank" rel="noopener">Lars Ruthotto&lt;/a>. In addition to this post, the team has also given a &lt;a href="">midterm presentation&lt;/a>, filmed a &lt;a href="https://youtu.be/8O9S8zm2N-E" target="_blank" rel="noopener">poster blitz video&lt;/a>, created a &lt;a href="http://www.math.emory.edu/site/cmds-reuret/resources/Image_Registration_Poster.png">poster&lt;/a> and written a &lt;a href="../../publications/watson-et-al-2024/">manuscript&lt;/a>.&lt;/p>
&lt;h2 id="image-registration">Image Registration&lt;/h2>
&lt;!-- ![](./resources/simple_IR.svg) -->
&lt;p>Image registration is the problem of finding a transformation which best &amp;ldquo;aligns&amp;rdquo; one image with another image. See &lt;a href="https://archive.siam.org/books/fa06/" target="_blank" rel="noopener">FAIR&lt;/a> for more detail on the problem. Some direct impacts of registration include&lt;/p>
&lt;ul>
&lt;li>
&lt;p>Medical Imaging: Accurate diagnosis, treatment planning, and monitoring of diseases.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Remote sensing: Analyze changes in the environment, such as deforestation, urbanization, and disaster assessment.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Computer Vision and Robotics: Object recognition, 3D reconstruction, and augmented reality applications.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="our-contributions">Our Contributions&lt;/h2>
&lt;p>Over the course of the eight-week REU program, our team successfully implemented a multitude of image registration tools into working code. The variety of tools we implemented showcases how machine learning frameworks like PyTorch and JAX simplify code and saves programming time. Particularily, we highlight how auto-differentiation makes programs simpler. Below, we outline three different projects we pursued during the program.&lt;/p>
&lt;h3 id="homotopy-optimization">Homotopy Optimization&lt;/h3>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="./resources/Affine_Homotopy.svg" alt="with and without homotopy" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;blockquote>
&lt;p>The far left column shows the template $\mathcal{T}$, to be aligned with the reference $\mathcal{R}$. The middle column shows a non-ideal solution, a local minimizer. The far right column shows how with homotopy, a global minimizer can be obtained.&lt;/p>&lt;/blockquote>
&lt;p>The main challenge in image registration is that the problem often lacks a clear or unique solution, and the approach to solving it is not straightforward. However, the problem becomes simpler when the images are blurred or smoothed. Computing the best transformation, or global minimizer, for a highly smoothed version of the image is more straightforward. Homotopy methods take advantage of this by first finding a global minimizer when the images are very smooth, and then tracing the path of the global minimizer through more and more detailed versions of the image back to the original problem. The figure above demonstrates that without homotopy optimization, we find a non-ideal transformation, but with homotopy omptimization, we can achieve a satisfactory one.&lt;/p>
&lt;p>The difficulty with implementing homotopy optimization for image registration lies in having to compute a Hessian matrix, which is a laborious process to derive. However, this is made simpler with automatic-differentiation.&lt;/p>
&lt;!-- ### Registration with Neural ODE
One approach that allows for greater freedom in addition to other nice properties is to construct the transformation as a flow field (see this [paper](https://pubmed.ncbi.nlm.nih.gov/29097881/)). This involves describing how the transformation changes over time, starting from a state that does nothing (the identity transformation), and using a velocity function to explain the changes. By observing the transformation at a specific time, in this case time 1, we get our final transformation. Neural networks, which are [great function approximators](https://en.wikipedia.org/wiki/Universal_approximation_theorem), can be used to approximate the velocity function. When this is done, it's referred to as a neural ordinary differential equation (neural ODE). The figure below illustrates registration with a neural ode and how it manages to create a complex and accurate registration.
![](./resources/NODE_hand-1.svg)
### Solving an Image Registration Problem
The problem of image registration can be phrased as minimizing an objective function that quantifies the error in the registration. This amounts to finding the parameters of the neural network velocity function that minimizes the objective. An effective method to avoid getting non-ideal solutions is by using a multi-scale approach (more [explanation](https://archive.siam.org/books/fa06/)). This involves solving the registration problem when the template and reference are "blurred", and then starting from that solution, solve successively less blurred versions of the problem until the blurring is gone. The figure below illustrates the multiscale approach.
![](./resources/theta_hands.svg)
#### Homotopy Optimization
Our team has explored homotopy optimization methods, experiencing both successes and challenges. Homotopy optimization aims to achieve what multiscale methods accomplish but operates on a continuous scale by tracing the path of a global minimizer. This technique leverages second-order optimization methods, starting with securing a minimizer for the blurry registration problem. However, integrating it effectively with neural ODE registration has proven to be difficult since hihgly blurred versions of the problem are still highly non-convex. -->
&lt;!--
## References
[1] A. Mang and L. Ruthotto. A lagrangian gauss–newton–krylov solver for mass-and intensity-
preserving diffeomorphic image registration. SIAM Journal on Scientific Computing,
39(5):B860–B885, 2017.
[2]
-->
&lt;h3 id="super-resolution">Super Resolution&lt;/h3>
&lt;p>Given a sequence of low resolution images of a subject in motion, the goal of super-resolution is to constructa high resolution of the first frame of the sequence. For example, a super resolution algorithm would be able to construct a high resolution image of a heart given only low resolution images taken at different points in time as the heart beats.&lt;/p>
&lt;p>The super-resolution problem is like the image registration problem in that the goal is to transformations that align images. Particularily, we find a transformation for each low resolution image that aligns it with the high resolution image. But since the high-res image is unkown, we construct it simultaneously, optimizing for both the transformations and the high resolution image together.&lt;/p>
&lt;figure>
&lt;img src="./resources/super-resolution-diagram.png"
alt="super-resolution diagram"
width=600>
&lt;figcaption>Given low resolution images $\mathbf{d}$, these each have a high resolution counterpart labeled by the letter $\mathbf{f}$, which we find through our optimization.
&lt;/figcaption>
&lt;/figure>
&lt;!--
Using Neural Ordinary Differential Equations ([Neural ODEs](https://arxiv.org/pdf/1806.07366)), we model the continuous transformation of the image over time, which allows for a single transformation function to describe the entire sequence.
-->
&lt;p>We employ advanced techniques like variable projection to enhance the optimization process, improving the accuracy and efficiency of super-resolution.&lt;/p>
&lt;h3 id="registration-with-intensity-dynamics">Registration with Intensity Dynamics&lt;/h3>
&lt;p>In medical imaging, a &lt;a href="https://www.ncbi.nlm.nih.gov/books/NBK557794/" target="_blank" rel="noopener">contrast agent&lt;/a> is often injected at the beginning of a sequence of images taken over time. This leads to observable &lt;em>dynamics&lt;/em> - intensity changes over time - which contain &lt;a href="https://link.springer.com/article/10.1007/s00330-003-2108-0" target="_blank" rel="noopener">medically relevant information&lt;/a>. Correctly observing these dynamics requires tracking the same region through each image, which makes image registration important for this problem. However, to solve a registration problem, we have relied on the assumption that intensity is preserved- that is, that no such dynamics are present (There are some &lt;a href="https://www.ams.org/journals/notices/202405/rnoti-p613.pdf" target="_blank" rel="noopener">interesting math&lt;/a> connections between intensity preservation and Neural ODEs).&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="./resources/dynamics_new.png" alt="Kidney MRI Images with Movement and Dynamics" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;blockquote>
&lt;p>The synthetic data in the above figure shows the data typical of this problem- the medulla and cortex of the kidneys experience different intensity changes over time after the injection of contrast. The original MRI is from Figure 4 of &lt;a href="https://doi.org/10.1007/978-3-642-31340-0_20" target="_blank" rel="noopener">Registration of Dynamic Contrast Enhanced MRI with Local Rigidity Constraint&lt;/a>.&lt;/p>&lt;/blockquote>
&lt;p>This problem is of the same mathematical form as the Super Resolution problem. &lt;a href="https://www.mic.uni-luebeck.de/fileadmin/mic/publications/2017/ediss1862.pdf" target="_blank" rel="noopener">Previous work&lt;/a> has used similar techniques as in super-resolution to estimate maps of parameters in various dynamical models. We are working on estimating the intensity changes directly.&lt;/p>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>In implementing three image-registration related problems with python machine learning tools, we have demonstrated the benefits of using these tools to further improve the state of image-registration. We have laid the groundwork for using homotopy methods for image registration, and have implemented variable projection with automatic differentiation. We hope that image registration researchers continue to take advantage of these tools in the future.&lt;/p></description></item><item><title>Efficient Processing of Image Sequences with Krylov Subspace Recycling</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2024-krylov/</link><pubDate>Mon, 08 Jul 2024 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2024-krylov/</guid><description>&lt;p>This blog post was written by Clara Armstrong, Olivia Kallay, and Srijon Sarkar and published with minor edits. The team was advised by &lt;a href="../author/lucas-onisk">Lucas Onisk&lt;/a>. In addition to this post, the team has also given a &lt;a href="">midterm presentation&lt;/a>, filmed a &lt;a href="">poster blitz video&lt;/a>, created a &lt;a href="">poster&lt;/a> and written a &lt;a href="">manuscript&lt;/a>.&lt;/p>
&lt;h2 id="background">Background&lt;/h2>
&lt;p>The image deblurring problem can be framed as a least-squares problem. We let $A$ be our blurring matrix, the convolution of a &lt;em>Point-Spread Function (PSF)&lt;/em> and a boundary condition. The true image is represented by $X$, and the blurred by $B$; we vectorize both to get our linear system, $Ax=b$. For example, we see below an image blurred using PSF Gauss with 2% noise:&lt;/p>
&lt;div style="display:flex; flex-direction:row;">
&lt;img src="captioned_blurred.png" alt="captioned_blurred" style="width:33%;">
&lt;img src="captioned_psf2.png" alt="captioned_psf2" style="width:33%;">
&lt;img src="captioned_deblurred.png" alt="captioned_deblurred" style="width:33%;">
&lt;/div>
&lt;p>Consider the solution $x$ written with respect to the singular value decomposition (SVD) assuming $A$ is invertible. The blurred imagine contains error such that $b=\hat{b}+e$; we observe that&lt;/p>
&lt;p>$$
x = V\Sigma^{-1}U^Tb
= \sum_{i=1}^{n}\frac{u_i^Tb}{\sigma_i}v_i
= \sum_{i=1}^{n}\frac{u_i^T\hat{b}}{\sigma_i}v_i+ \sum_{i=1}^{n}\frac{u_i^Te}{\sigma_i}v_i
$$&lt;/p>
&lt;p>where the first sum is our &lt;em>true solution&lt;/em> and the second is our &lt;em>inverted noise&lt;/em>. For image deblurring problems, the $u_i^Te$ terms are constant in magnitude. $A$ is an ill-conditioned matrix, causing the solution to be dominated by the inverted noise, which becomes very big as the singular values $\sigma_i$ decay to numerical zero.&lt;/p>
&lt;h2 id="our-approach">Our Approach&lt;/h2>
&lt;p>We will focus on two Krylov techniques: &lt;em>Arnoldi&lt;/em> and &lt;em>Golub-Kahan bidiagonalization&lt;/em>. Using Krylov methods allows us to work in a smaller subspace that captures enough information about the system to compute an approximate solution, making our problem easier to solve than solving directly.&lt;/p>
&lt;p>A Krylov subspace is the span of repeated applications of a matrix $A$ to a vector $b$, or&lt;/p>
&lt;p align="center">$K_p(A, b) = \text{span}\{b, Ab, A^2b, ... , A^{p-1}b\}$&lt;/p>
&lt;p>The Arnoldi relation is at the core of the &lt;em>Generalized Minimal Residual (GMRES)&lt;/em> method for solving &lt;strong>square nonsymmetric&lt;/strong> linear problems. After $p$ steps of the Arnoldi iteration, we develop the relation&lt;/p>
&lt;p align="center">$AV_p = V_{p+1}H_{p+1}$&lt;/p>
&lt;p>where $A$ is the square operator from the linear system, $V_{p+1}$ has orthonormal columns that span the subspace $K_p(A, b)$, $V_p$ contains the first $p$ orthonormal columns of $V_{p+1}$, and $H_{p+1}$ is an upper Hessenberg matrix (an upper triangular matrix that contains an additional subdiagonal band) of the scalar coefficients from the orthonormalization process of Gram-Schmidt, used in Arnoldi.&lt;/p>
&lt;p>&lt;em>Golub-Kahan bidiagonalization (GKB)&lt;/em> is associated with the LSQR method, where $A$ could be &lt;strong>non-square&lt;/strong>. GKB generates orthonormal vectors that span the following spaces:&lt;/p>
&lt;p>$$K_p(A^TA, A^Tb) \text{ and } K_p(AA^T, b)$$&lt;/p>
&lt;p>After $p$ steps of GKB, we have the following relationships&lt;/p>
&lt;p>$$AV_p = U_{p+1}B_{p+1} \text{ and } A^TU_{p+1} = V_{p+1}\tilde B_{p+1}$$&lt;/p>
&lt;p>where $V_{p+1}$ and $U_{p+1}$ contain orthonormal columns spanning $K_p (A^TA, A^Tb)$ and $K_p(AA^T, b)$, $B_{p+1}$ is the lower-bidiagonal, projected variant of $A$, and $\tilde B_{p+1}$ is square due to the inclusion of an additional column.&lt;/p>
&lt;h2 id="our-observations">Our Observations&lt;/h2>
&lt;p>Since GMRES and LSQR require many matrix-vector products, we desire to gain efficiency through recycling strategies to circumvent such repetitive computation while solving sequences of closely related problems.&lt;/p>
&lt;p>We observed for that $r^{(2)}=Ax_{approx}^{(1)}-b^{(2)}$, the residual of the second problem using our solution to the first (seed) problem, the Reconstructive Residual Norm (RRN) satisfies&lt;/p>
&lt;p>$$DP \leq \frac{||r^{(2)}||}{||b||} &amp;lt; 1,$$&lt;/p>
&lt;p>where DP represents a breakout parameter from the &lt;em>Discrepancy Principle&lt;/em>, $\tau\delta$, in which $\tau$ is a safety parameter close to one and $\delta$ is an upper-bound on the erroneous right-hand sides.&lt;/p>
&lt;p align="center">&lt;img width="334" alt="captioned_blurred" src="recycling.jpg" &lt;/p>
&lt;p>Although the residual is above the breakout level given by the DP, it remains significantly lower than if we started from the zero vector (which would produce a relative residual of one). Thus, by recycling basis vectors from previous solutions to solve subsequent problems, we drastically reduce the RRN.&lt;/p>
&lt;p>We note that with these types of methods, residuals decrease, but errors demonstrate semiconvergent behavior; we see the error decrease until a certain point at which it will begin growing exponentially.&lt;/p>
&lt;p>As the basis from the seed problem contains useful information for non-arbitrarily chosen right-hand sides (image frames for our purposes), we incorporate this information by calculating&lt;/p>
&lt;p>$$\text{res} = \frac{r^{(2)} - V_pV_p^Tr^{(2)}}{||r^{(2)} - V_pV_p^Tr^{(2)}||}$$&lt;/p>
&lt;p>and appending it in $[V_p \quad res] = \tilde V_{\ell}$ thereby forming the &lt;em>flexible Arnoldi&lt;/em> method as $A[\tilde V_{\ell}] = \tilde U_{\ell+1}J_{\ell+1}$ at step $\ell$. $U$ and $J$ come from completing the back half of Arnoldi, respectively representing a new basis and a projected variant of $A$.&lt;/p>
&lt;h2 id="future-directions">Future Directions&lt;/h2>
&lt;p>Eventually, we will run out of memory as we continue iterating flexible Arnoldi and storing more basis vectors. To avoid this, we will utilize different compression strategies to decide which basis vectors should be retained for solving subsequent problems, and which may be discarded. The compression strategies we aim to use are&lt;/p>
&lt;ol>
&lt;li>Truncated Singular Value Decomposition&lt;/li>
&lt;li>Reduced Basis Decomposition&lt;/li>
&lt;li>Solution-Oriented Compression&lt;/li>
&lt;/ol>
&lt;p>We aim to combine compression techniques with flexible algorithms, including extending our work to non-square operators to replicate &lt;em>flexible GKB&lt;/em> for LSQR.&lt;/p>
&lt;p>Further, as these iterative processes continue and our subspace grows, it might capture information from small singular values. So, we discern the need for penalized least squares using &lt;em>Tikhonov regularization&lt;/em> for our projected problems as&lt;/p>
&lt;p>$${\min_\textbf{x}} ||A\textbf{x}-\textbf{b}||_2^2 + \lambda||\textbf{x}||_2^2.$$&lt;/p>
&lt;p>So far, we have only focused on sequences of images, such as those that may arise from a video. However, these techniques can also be useful for other sequences of linear inverse problems, such as those arising from medical imaging.&lt;/p>
&lt;h2 id="references">References&lt;/h2>
&lt;ol>
&lt;li>&lt;em>On the Use of Arnoldi and Golub-Kahan Bases to Solve Nonsymmetric Ill-Posed Inverse Problems&lt;/em>, Brown, A. M., (2015)&lt;/li>
&lt;li>&lt;em>GMRES: A Generalized Minimal Residual Algorithm for Solving Nonsymmetric Linear Systems&lt;/em>, Saad, Y. and Schultz, M.H., (1986)&lt;/li>
&lt;li>&lt;em>Hybrid Projection Methods with Recycling for Inverse Problems&lt;/em>, Jiang, J., Chung, J., and de Sturler, E., (2021)&lt;/li>
&lt;/ol></description></item><item><title>Improving VAEs with Normalizing Flows</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2024-vae/</link><pubDate>Mon, 01 Jul 2024 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2024-vae/</guid><description>&lt;p>This post was written by &lt;a href="https://www.linkedin.com/in/cbertley/" target="_blank" rel="noopener">Callihan Bertley&lt;/a>, &lt;a href="https://www.linkedin.com/in/claire-gan-758630293/" target="_blank" rel="noopener">Claire Gan&lt;/a>, &lt;a href="https://www.linkedin.com/in/rishi-leburu-751430298/" target="_blank" rel="noopener">Rishi Leburu&lt;/a>, and &lt;a href="https://www.linkedin.com/in/maliawalewski/" target="_blank" rel="noopener">Malia Walewski&lt;/a>. The team was advised by Dr. Deepanshu Verma. In addition to this post, we have also created &lt;a href="./Team-VAE-Midterm-pres.pdf">slides&lt;/a> for a midterm presentation, a &lt;a href="https://youtu.be/qmQ--692cvc" target="_blank" rel="noopener">poster blitz video&lt;/a>, and a &lt;a href="./Poster.pdf">poster&lt;/a>.&lt;/p>
&lt;h3 id="project-overview">Project Overview:&lt;/h3>
&lt;p>Variational Autoencoders (VAEs) have emerged as powerful deep generative models in recent years. Our project aims to improve the expressiveness of VAEs by incorporating different types of normalizing flows, specifically Inverse Autoregressive Flows (IAFs) and Partially Convex Potential Maps (PCP-Maps).&lt;/p>
&lt;h4 id="what-are-vaes">What are VAEs?&lt;/h4>
&lt;p>VAEs use neural network architectures to learn a compact representation of input data through an encoder and generate new data through a decoder. The encoder compresses the input into a latent vector, outputting a mean and standard deviation for each latent variable to approximate the true latent representation (posterior distribution). Samples from this distribution should resemble the input when passed through the decoder.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="./VAE_2.png" alt="VAE pic" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;em>Figure 1: Architecture of Variational Autoencoder visualization&lt;/em>&lt;/p>
&lt;p>VAEs typically use a Gaussian approximation for the posterior distribution.
&lt;strong>However, since the true posterior distribution often deviates from normality, we aim to find a more expressive posterior.&lt;/strong>&lt;/p>
&lt;h4 id="how-to-improve-vaes">How to improve VAEs?&lt;/h4>
&lt;p>To improve VAEs, we can apply normalizing flows (NFs) to the Gaussian approximated posterior distribution from the encoder. NFs apply a series of invertible transformations, parameterized by neural networks that can be optimized during VAE training.&lt;/p>
&lt;p>The figure below demonstrate how normalizing flow are able to transform into a simple distribution to a more complex distribution.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="./NF.jpg" alt="My Photo" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;em>Figure 2: Comparison between before (left) and after (right) of training using NFs.&lt;/em>&lt;/p>
&lt;p>Our goal is to show that models with normalizing flows can perform better than standard VAEs.&lt;/p>
&lt;h3 id="methods">Methods&lt;/h3>
&lt;p>To address this challenge, we compare three different models with VAE using the MNIST handwriting dataset. In our experiment, we follow a similar implementation to &lt;a href="https://github.com/lollcat/Autoencoders-deep-dive/blob/Pytorch/Report.pdf" target="_blank" rel="noopener">L. Midgley&amp;rsquo;s&lt;/a> standard VAE and IAF VAE. Then we implemented the PCP-Map.&lt;/p>
&lt;ol>
&lt;li>
&lt;p>&lt;strong>Standard VAE&lt;/strong>
The standard VAE uses a Gaussian approximation for both the encoder and latent distribution. For further details, see &lt;a href="https://arxiv.org/abs/1312.6114" target="_blank" rel="noopener">1&lt;/a> and &lt;a href="https://arxiv.org/abs/2103.05180" target="_blank" rel="noopener">4&lt;/a>.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Inverse Autoregressive Flow (IAF) VAE&lt;/strong>
The IAF VAE applies a series of invertible, autoregressive transformations to the Gaussian approximated distribution in the encoder. For each transformation, an autoregressive neural network takes in the input latent sample and a conditional vector dependent on the given inputs and outputs mean and standard deviation of the tranfromation.&lt;/p>
&lt;p>The Jacobian matrices of the transformation are lower triangular, which makes computing the determinant, and hence the loss, not expensive.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="./IAF_architecture.png" alt="IAF pic" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;em>Figure 3: Architecture of IAF VAE visualization&lt;/em>&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Partially Convex Potential Map&lt;/strong>
The PCP-Map VAE is another way to transform the input, Gaussian distribution to a more complex distribution. This is achieved by parameterizing a transformation as the gradient of a scalar-valued partially input convex neural network. These maps are constructed to have a triangular Jacobian matrix. This structure allows the Jacobian to be computed as the Hessian of the PICNN.&lt;/p>
&lt;/li>
&lt;/ol>
&lt;p>&lt;strong>Optimizing Evidence Lower Bound (ELBO)&lt;/strong>
We use the KL-Divergence loss function to measure model performance, aiming to minimize the KL divergence between our approximate posterior and the true posterior by minimizing the negative Evidence Lower Bound (ELBO). The figure belwo shows how ELBO is measured.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="./ELBO.png" alt="ELBO pic" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;em>Figure 4: Understanding of the KL-Divergence (Image Source: &lt;a href="https://mbste.github.io/posts/vae/" target="_blank" rel="noopener">https://mbste.github.io/posts/vae/&lt;/a>)&lt;/em>&lt;/p>
&lt;p>We compare the models using negative ELBO loss, reconstructed images, and contour plots.&lt;/p>
&lt;h3 id="results-and-conclusion">Results and Conclusion&lt;/h3>
&lt;p>After training for approximately 2000 epochs, the IAF VAE produced slightly better reconstructed images than the standard VAE. The contour plot for the IAF VAE showed a non-Gaussian distribution more similar to the input data distribution, suggesting improved expressiveness.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th style="text-align: center">
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Image 1" srcset="
/site/cmds-reuret/projects/2024-vae/_test_imgs_hu_6c6e9fe75b14419c.webp 400w,
/site/cmds-reuret/projects/2024-vae/_test_imgs_hu_fc0dbae35dae4763.webp 760w,
/site/cmds-reuret/projects/2024-vae/_test_imgs_hu_b8b6c8cc5e4286f9.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2024-vae/_test_imgs_hu_6c6e9fe75b14419c.webp"
width="758"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/th>
&lt;th style="text-align: center">
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Image 2" srcset="
/site/cmds-reuret/projects/2024-vae/full_std_reconstruction_imgs_hu_d35156e162964a7.webp 400w,
/site/cmds-reuret/projects/2024-vae/full_std_reconstruction_imgs_hu_6ee2fb130ad15217.webp 760w,
/site/cmds-reuret/projects/2024-vae/full_std_reconstruction_imgs_hu_3ed5dd414a3f6c7f.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2024-vae/full_std_reconstruction_imgs_hu_d35156e162964a7.webp"
width="758"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/th>
&lt;th style="text-align: center">
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Image 3" srcset="
/site/cmds-reuret/projects/2024-vae/full_iaf_reconstruction_imgs_hu_5b91d4bd51c83581.webp 400w,
/site/cmds-reuret/projects/2024-vae/full_iaf_reconstruction_imgs_hu_926b0d9f6dedd11b.webp 760w,
/site/cmds-reuret/projects/2024-vae/full_iaf_reconstruction_imgs_hu_faa961ee12d8a882.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2024-vae/full_iaf_reconstruction_imgs_hu_5b91d4bd51c83581.webp"
width="752"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td style="text-align: center">
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Image 4" srcset="
/site/cmds-reuret/projects/2024-vae/data_hu_832ea0e276a9ab90.webp 400w,
/site/cmds-reuret/projects/2024-vae/data_hu_5bbdc399308b8eb0.webp 760w,
/site/cmds-reuret/projects/2024-vae/data_hu_e118a81b62ded3ab.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2024-vae/data_hu_832ea0e276a9ab90.webp"
width="760"
height="575"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/td>
&lt;td style="text-align: center">
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Image 5" srcset="
/site/cmds-reuret/projects/2024-vae/std_hu_f1251fdebb31afcf.webp 400w,
/site/cmds-reuret/projects/2024-vae/std_hu_cf53b93d3328cf2b.webp 760w,
/site/cmds-reuret/projects/2024-vae/std_hu_973534e169b3ef9.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2024-vae/std_hu_f1251fdebb31afcf.webp"
width="760"
height="580"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/td>
&lt;td style="text-align: center">
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Image 6" srcset="
/site/cmds-reuret/projects/2024-vae/iaf_hu_8fc7f67918fba53b.webp 400w,
/site/cmds-reuret/projects/2024-vae/iaf_hu_95096e8647b0e014.webp 760w,
/site/cmds-reuret/projects/2024-vae/iaf_hu_183df23f5f0437b8.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2024-vae/iaf_hu_8fc7f67918fba53b.webp"
width="760"
height="577"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;h2 id="references">References&lt;/h2>
&lt;p>[1] D. P. Kingma and M. Welling, &amp;ldquo;Auto-Encoding Variational Bayes,&amp;rdquo; 2014.&lt;/p>
&lt;p>[2] D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, and M. Welling, &amp;ldquo;Improving variational inference with inverse autoregressive flow,&amp;rdquo; 2017&lt;/p>
&lt;p>[3] L. Midgley, &amp;ldquo;Improving variational inference with inverse autoregressive flow,&amp;rdquo; University of Cambridge, Tech. Rep. 2021. [Online]. Available: &lt;a href="https://github.com/lollcat/Autoencoders-deep-dive/blob/Pytorch/Report.pdf" target="_blank" rel="noopener">https://github.com/lollcat/Autoencoders-deep-dive/blob/Pytorch/Report.pdf&lt;/a>&lt;/p>
&lt;p>[4] L. Ruthotto and E. Haber, &amp;ldquo;An introduction to deep generative modeling,&amp;rdquo; GAMM-Mitt., vol. 44, no. 2, pp. Paper No.e202 100 008, 24, 2021. [Online]. Available: &lt;a href="https://doi.org/10.1002/gamm.202100008" target="_blank" rel="noopener">https://doi.org/10.1002/gamm.202100008&lt;/a>&lt;/p>
&lt;p>[5] Z. O. Wang, R. Baptista, Y. Marzouk, L. Ruthotto, and D. Verma, &amp;ldquo;Efficient neural network approaches for conditional optimal transport with applications in bayesian inference,&amp;rdquo; 2023. [Online]. Available: &lt;a href="https://arxiv.org/abs/2310.16975" target="_blank" rel="noopener">https://arxiv.org/abs/2310.16975&lt;/a>&lt;/p></description></item><item><title>Optimal Experiment Design and Image Reconstruction using Generative Methods</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2024-oed/</link><pubDate>Mon, 01 Jul 2024 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2024-oed/</guid><description>&lt;p>This blog post was written by Spiros Manolas, Anish Mitagar, and Nela Riddle and published with minor edits. The team was advised by &lt;a href="../author/nicole-yang">Dr. Nicole Yang&lt;/a>.In addition to this post, the team has also given a &lt;a href="">midterm presentation&lt;/a>, filmed a &lt;a href="">poster blitz video&lt;/a>, created a &lt;a href="">poster&lt;/a> and written a &lt;a href="">manuscript&lt;/a>.&lt;/p>
&lt;div style="text-align: center;">
&lt;img src="img_assets/Problem.png" alt="Sample Image" width="75%">
&lt;/div>
&lt;br>&lt;br>
Reconstructing images from noisy and indirect observations is an ill-posed inverse problem critical for many medical and imaging applications However, obtaining such indirect observations is expensive time and cost wise. We wish to improve the measurement process by finding the best sub-sampled indirect measurements/observations to take, and reconstruct a quality image from the best sub-sampled set.
&lt;div style="text-align: center;">
&lt;h3>Example Training Process&lt;/h3>
&lt;/div>
&lt;div style="text-align: center;">
&lt;img src="img_assets/TrainModel.png" alt="Sample Image" width="90%">
&lt;/div>
&lt;div style="text-align: center;">
&lt;h3>Example Inference Process&lt;/h3>
&lt;/div>
&lt;div style="text-align: center;">
&lt;img src="img_assets/InferenceModel.png" alt="Sample Image" width="90%">
&lt;/div>
&lt;br>&lt;br>
We explore how Normalizing Flows, a generative method, and autoencoders can be leveraged to learn quality image generation from a sub-sampled set of indirect observations. We also explore introducing a differentiable or learnable mask function into our model as design parameter d, to potentially learn the best sub-sampled set of indirect observations to reconstruct from using Normalizing Flows.</description></item><item><title>Accurately Classifying Out-Of-Distribution Data in Facial Recognition</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2023-ood/</link><pubDate>Fri, 01 Dec 2023 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2023-ood/</guid><description>&lt;!-- https://docs.google.com/document/d/1u7_vNpwToqwOdBlpbuRNXYDAMrgsaymMO8iSxA8UqNw/edit -->
&lt;p>This blog post was written by Gianluca Barone, Aashirit Cunchala, and Rudy Nunez and published with minor edits. The team was advised by &lt;a href="https://nicoletyang.github.io/" target="_blank" rel="noopener">Nicole Yang&lt;/a>. In addition to this post, the team gave a midterm presentation, made a poster blitz video, wrote a paper, and won the best poster award for their &lt;a href="content/2023-OOD-Poster.pdf">poster&lt;/a>.&lt;/p>
&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>As facial recognition technology is deployed at large scales, it is crucial to ensure its fairness. A fair model has approximately equal performance in all different classes. However, models are often unfair due to bias, which can occur for several reasons, such as having an unbalanced training set. One common assumption in machine learning is that the training and test images are from the same distribution, but in real-world applications, that is often not the case; for example, if a model is trained on an imbalanced training set, it may not do well on a balanced set. Due to this, the minority class can be overpowered by the majority class.&lt;/p>
&lt;p>The core goal of this project is to increase a model’s fairness when encountering images outside of the training distribution. Such
out-of-distribution (OOD) images belong to a distribution different from the one the model was trained on. We used two different face datasets to have different distributions: our training dataset was UTKFace, an imbalanced dataset of 20,000 images, and our testing dataset was FairFace, a balanced dataset of 100,000 images.&lt;/p>
&lt;h2 id="outlier-exposure">Outlier Exposure&lt;/h2>
&lt;p>One approach to making a model more familiar with out-of-distribution data is through &lt;a href="https://arxiv.org/abs/1812.04606" target="_blank" rel="noopener">outlier exposure&lt;/a>. In outlier exposure, a model is exposed to both in-distribution and out-of-distribution data during training stages. When it encounters the OOD data during testing, it is more likely to classify it accurately. To do this, we attempted to collect the furthest OOD images from each dataset. One approach to finding these images was using Kullback–Leibler (KL) divergence based on pixels. KL divergence measures the distance between two different distributions. Using this, we found the images with the largest KL distance, which means they are the furthest from the distribution.&lt;/p>
&lt;p>In outlier exposure, the loss function is modified to include an additional term. In this term, there is a parameter lambda that scales the significance of the other term, which we were able to train to find an optimal lambda. We used the KL divergence to relate the loss function to the distance between distributions.&lt;/p>
&lt;p>Using the KL score, we have noticed that UTKFace and Fairface have quite similar distributions due to the low KL score. Due to this, outlier exposure may not be as effective since it does better when the distributions are more dissimilar. To rectify this, we increased the training set by combining UTKFace with 20% of outliers from FairFace. We found these outliers by collecting the 20% of images that are the furthest from the mean distribution of FairFace chosen through KL divergence.&lt;/p>
&lt;h2 id="main-results">Main Results&lt;/h2>
&lt;p>We observe that the accuracy and other metrics of the model can be improved by applying Outlier Exposure, incorporating a trainable weight parameter to increase the model&amp;rsquo;s emphasis on outlier images, and by re-weighting the importance of different class labels.
Some items of future work include using feature norms of the images and other outputs of the CNN to sort the
images and improve classification. Further, one may adapt the activation features code into the outlier exposure code to see if it is
more accurate than using the KL divergence based on pixels.&lt;/p>
&lt;h2 id="reflection">Reflection&lt;/h2>
&lt;p>This project taught us much about neural networks and how to improve facial recognition models. We also learned more about writing manuscripts and doing research. We plan on continuing our work on this project and exploring the possibility of using image features and additional weight parameters to classify out-of-distribution data.&lt;/p>
&lt;h2 id="interested-in-learning-more">Interested in Learning More?&lt;/h2>
&lt;p>Please see our award-winning &lt;a href="content/2023-OOD-Poster.pdf">poster&lt;/a>.&lt;/p></description></item><item><title>Fast &amp; Fair: Efficient Training of Fair Neural Networks</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2023-fast-and-fair/</link><pubDate>Fri, 01 Dec 2023 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2023-fast-and-fair/</guid><description>&lt;!-- https://docs.google.com/document/d/1aIJEDrfuTmHEsgWYf9xfC-uX9ub8IPgOiKTmKlBMgck/edit#heading=h.dou3952sp91t -->
&lt;p>This blog post was written by Allen Minch, Hung Anh Vu, and Annie Warren and published after some minor edits. The team was advised by &lt;a href="../author/elizabeth-newman/">Dr. Elizabeth Newman&lt;/a>. More information about this project can be found in our paper, &lt;a href="https://youtu.be/7J0WmeorOF8" target="_blank" rel="noopener">poster blitz video&lt;/a>, and &lt;a href="content/2023-Fast-And-Fair-Poster.pdf">poster&lt;/a>.&lt;/p>
&lt;h2 id="unfairness-in-machine-learning">(Un)Fairness in Machine Learning&lt;/h2>
&lt;p>Our project focuses on a key problem in machine learning: unfairness. How do we define unfairness, and where does it come from?&lt;/p>
&lt;p>Classifiers in machine learning, such as deep neural networks, can be very good at classifying complex data, such as images. However, machine learning classifiers trained through supervised learning are designed to minimize the average classification error in the data they are trained on. This means they are inherently designed to be very “faithful” to the data they are trained on and to reflect patterns and correlations in that data. This can be a problem when such correlations involve a sensitive attribute - for instance, a correlation of a person’s race with an attribute we believe should not be related to race.&lt;/p>
&lt;p>Suppose the real-world data set that a machine learning classifier is trained on is biased with respect to some sensitive attribute, like race. In that case, the classifier will, in trying to maximize its accuracy on the data it is being trained on, be “faithful” to the bias that it is given and be trained to make predictions that reflect this bias. While it could be fairly stated that these biased predictions are not the fault of the machine learning algorithm itself but the data on which it is trained, that doesn’t change the fact that, as a society, we would generally view it as unfair to base real-world decisions on using a machine learning classifier that makes biased predictions. This is especially true when the stakes are high, where a biased classifier prediction could mean someone getting unfairly arrested, unfairly not being hired for a job, or unfairly being denied a loan. It is thus in our interests as a society to be careful about using machine learning classifiers that might be biased in their predictions and try to tweak how they are trained to reduce this unfairness.&lt;/p>
&lt;h2 id="improving-fairness-through-adversarial-training">Improving Fairness Through Adversarial Training&lt;/h2>
&lt;p>We develop efficient adversarial training techniques to reduce the unfairness of classifiers in machine learning in the context of a sensitive attribute like race. In our project, we are dealing with supervised machine learning, where a classifier is trained on a dataset with many samples of input features and a corresponding true label. Because the context is supervised, we have true labels as information to help us evaluate the classifier&amp;rsquo;s performance. To assess our success in improving the fairness of a classifier, we need to have fairness metrics by which to evaluate a classifier.&lt;/p>
&lt;p>Why would we think that adversarial training would help improve fairness? There are a couple of lenses for understanding why we might expect this to improve fairness. One lens has to do with the idea of overfitting. Intuitively, a primary reason we expect a classifier to be unfair is that it is overfitting to the bias in the particular data it is trained on. Adversarial training, by nature, tries to combat overfitting to the training data with its efforts to make a classifier more robust by giving perturbations to the training data points. Thus, we might expect combating overfitting with adversarial training also to combat unfairness, making the classifier less sensitive to bias in the given data. A second lens is that intuitively, a situation we would tend to view as most flagrantly unfair is when two individuals are quite similar yet happen to be part of different sensitive groups and turn out to be classified differently by a classifier. Perhaps some unfairness we observe in a classifier manifests in this kind of scenario. If we can make our classifier more robust - so that two nearby points are more likely to be classified the same way - perhaps this kind of unfairness could be reduced, which could minimize the unfairness of our classifier overall. Since adversarial training improves robustness, this would be another lens to understand why we might expect adversarial training to improve fairness.&lt;/p>
&lt;p>We develop a second-order optimization method for solving the adversarial training problem. When solving an optimization problem, first-order methods are used when we only need to compute the gradient. This means we rely solely on first-order information to optimize our function. In the unconstrained case, we can take a step as large as necessary in the direction of the gradient for maximal increase/decrease. However, for our approach, we must satisfy a constraint requiring us to project our solution onto the constraint set. Computing the gradient and projecting the solution can become very computationally expensive. Therefore, with some mathematical speculation, we could utilize second-order information (computing the Hessian) to solve the optimization problem more efficiently.&lt;/p>
&lt;h2 id="evaluating-fairness">Evaluating Fairness&lt;/h2>
&lt;p>We used multiple metrics for evaluating fairness with respect to sensitive attributes. To describe these intuitively, imagine that we had a dataset of individuals whose race is either white or non-white and with true labels as to whether or not they have committed a crime in the past. Suppose we used this data to train a classifier designed to predict whether or not an individual will commit a crime in the future and then used a test dataset to see how the classifier performs.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Independence&lt;/strong> This means that the prediction of our classifier is uncorrelated with the sensitive attribute; a positive prediction is equally likely for all groups. In our example, independence being satisfied in a test dataset would mean that if 15% of white individuals in the dataset are predicted to commit a crime in the future, then 15% of non-white individuals are also. (Note: Y represents the true label, Y_hat the predicted label, and S the sensitive attribute in the graphic below. The graphic shows two probabilities that would have to be equal for independence to be satisfied).&lt;/li>
&lt;li>&lt;strong>Separation&lt;/strong> This means that the prediction of our classifier is conditionally independent of the sensitive attribute, given an individual’s true label.&lt;/li>
&lt;li>&lt;strong>Sufficiency&lt;/strong>: This means that our classifier’s predictive usefulness is equal across sensitive groups. If sufficiency is satisfied, then for each $(x, y)$, the probability that an individual’s true label is $x$ given the value of their predicted label is $y$ is the same across sensitive groups.&lt;/li>
&lt;/ul>
&lt;p>We implemented these fairness metrics in Python with PyTorch tensors. We worked with binary classification and sensitive attributes in our project, which meant that being given PyTorch tensors containing each individual’s sensitive attribute, true label, and predicted label respectively, we could obtain the proportions needed for evaluating our fairness metrics leveraging some boolean indexing tricks.&lt;/p>
&lt;h2 id="coding-strategies-and-main-findings">Coding Strategies and Main Findings&lt;/h2>
&lt;p>Our adversarial training problem consists of two optimization problems: an inner one and an outer one. The outer optimization problem is the typical machine learning problem of finding optimal classifier parameters. The inner optimization problem consists of, for each point in the training data set, finding a perturbation to that point within some radius that maximizes the value of the loss function. In practice, to computationally solve a two-layered optimization problem like this, the approach we take is to solve the inner optimization problem at each point, update parameters in the outer optimization problem based on these solutions, solve the inner optimization problem again, and continue going back and forth between the two problems.&lt;/p>
&lt;p>As it turns out, robust training, because it requires solving the inner optimization problem at every training data point in every epoch of the outer optimization problem, can be pretty slow computationally compared to simply doing non-robust training. Thus, as the FastNFair team, we focused on solving the inner optimization problem as quickly and efficiently as possible.&lt;/p>
&lt;p>Our numerical results on three datasets show that our method is more efficient than first-order methods.
We also observe that adversarial training can improve fairness but potentially at the cost of accuracy.
If robust training with a certain radius improves fairness, it appears to
improve fairness by larger margins than random perturbation; solving
the optimization problem well is worthwhile.&lt;/p>
&lt;h2 id="interested-in-learning-more">Interested in Learning More?&lt;/h2>
&lt;p>Please see our &lt;a href="content/2023-Fast-And-Fair-Poster.pdf">poster&lt;/a>.&lt;/p></description></item><item><title>Investigating Impacts of Environmental and Socioeconomic Data</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2023-no2/</link><pubDate>Fri, 01 Dec 2023 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2023-no2/</guid><description>&lt;!-- https://docs.google.com/document/d/1DbfqmpkV7B7h78s44ZPp59s8fOqnl93ATvakM_nx7o8/edit -->
&lt;p>This blog post was written by Riley Chen, Mason Lu, Matilda Slosser, and Aneesh Srinivas and published with minor edits. The team was advised by &lt;a href="http://www.math.emory.edu/~jmchung/" target="_blank" rel="noopener">Julianne Chung&lt;/a>, &lt;a href="http://www.math.emory.edu/~mchun45/" target="_blank" rel="noopener">Matthias Chung&lt;/a>, &lt;a href="../author/elizabeth-newman/">Elizabeth Newman&lt;/a>. In addition to this post, the team has also created a &lt;a href="content/2023-NO2-Poster.pdf">poster&lt;/a>.&lt;/p>
&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>Nitrogen Dioxide ($NO_2$) is a common gaseous pollutant. It is a by-product of combustion, making its local concentration levels directly tied to gas vehicle emissions and fossil fuel-sourced energy production. In the United States, cars are the largest source of $NO_2$, which links high concentrations to urban environments.&lt;/p>
&lt;p>Tracking and predicting $NO_2$ concentrations is an issue of public health and social justice. Certain individuals, like children, older adults, or those with pre-existing lung conditions, are susceptible to poor health outcomes after exposure to $NO_2$. Marginalized people tend to be forced through economic and social structures into areas with relatively high $NO_2$ concentrations, resulting in disproportionate exposure.&lt;/p>
&lt;p>&lt;a href="https://pubmed.ncbi.nlm.nih.gov/31851499/" target="_blank" rel="noopener">Our data&lt;/a> is generated from the Air Quality System (AQS) $NO_2$ monitoring networks over the contiguous United States of the Environmental Protection Agency (EPA). This data is from Jan 1st, 2000, to Dec 31st, 2016, with 6210 points. Our main goal is to create a model to predict average $NO_2$ concentration across the contiguous United States. We also aimed to analyze the relationship between $NO_2$ concentration and the Social Vulnerability Index (SVI) geographically.&lt;/p>
&lt;h2 id="model-driven-approach">Model-Driven Approach&lt;/h2>
&lt;p>We plotted the daily average concentration of $NO_2$ over the US from 2000 to 2016 and observed that the data has seasonal oscillations and a decaying trend.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="NO2 Data" srcset="
/site/cmds-reuret/projects/2023-no2/images/data_hu_c1c55668e06d09ed.webp 400w,
/site/cmds-reuret/projects/2023-no2/images/data_hu_3358c2b2bd822385.webp 760w,
/site/cmds-reuret/projects/2023-no2/images/data_hu_63853e4276557a31.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2023-no2/images/data_hu_c1c55668e06d09ed.webp"
width="397"
height="327"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Therefore, we incorporated this information to build the model shown below.&lt;/p>
&lt;p>$ y_{\mathrm {model}}=p_1+p_2 e^{p_3(t-p4)}+p_5 \cos(2πp_6(t-p_7)). $&lt;/p>
&lt;p>where&lt;/p>
&lt;ul>
&lt;li>$p_1$ - our intercept (average $NO_2$ on December 31st, 1999).&lt;/li>
&lt;li>$p_2$ - starting value of the exponential decay&lt;/li>
&lt;li>$p_3$ - exponential decay scale factor&lt;/li>
&lt;li>$p_4$ - the starting time. Since the data started in 2000, we expected $p_4 ≈2000$.&lt;/li>
&lt;li>$p_5$ - the amplitude of the seasonal oscillation.&lt;/li>
&lt;li>$p_6$ - the frequency of the cosine component.&lt;/li>
&lt;li>$p_7$ - the cosine shift, which increases our model’s adaptability&lt;/li>
&lt;/ul>
&lt;p>To learn more about the parameters, we aimed to randomly generate samples, hoping that if the sample space was large enough, a good approximation for the distribution of parameters could be found. A Markov Chain Monte Carlo (MCMC) method was utilized to generate the samples, which implemented Bayes&amp;rsquo; theorem. The specific algorithm we used is called Adaptive Metropolis. We fixed $p_4$ to be 2000 and $p_6$ to be 1 to ensure enough random walk for the algorithm. We generated 400,000 sets of values for the parameter set.&lt;/p>
&lt;p>Plotting the samples component-wise, we observe that each graph only has one peak. That means that the samples agree and gives us some confidence that each parameter is within the range in its own graph. After investigating each parameter, we wanted to explore how each pair of parameters was related. Most were Gaussian if the outliers were ignored. However, $p_1$ and $p_3$ were highly correlated.&lt;/p>
&lt;p>We also calculated the posterior function value of the new parameter set for each random set we generated. We then updated our current maximize-a-posterior (MAP) estimate for the parameter set. This estimate gives us the most probable value for the parameter sets given this data. The model with this parameter set is plotted with a 95 percent prediction interval, as shown below. The interval is wider at the peaks and troughs, making it harder to predict the exact value at the highest and lowest points in each oscillation cycle.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="MAP Estimate" srcset="
/site/cmds-reuret/projects/2023-no2/images/fit_hu_f9ce4e15ca21002c.webp 400w,
/site/cmds-reuret/projects/2023-no2/images/fit_hu_9570cc3d00140d6d.webp 760w,
/site/cmds-reuret/projects/2023-no2/images/fit_hu_1d34f047088656ea.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2023-no2/images/fit_hu_f9ce4e15ca21002c.webp"
width="518"
height="404"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h2 id="main-findings">Main Findings&lt;/h2>
&lt;p>Our model-driven approach with appropriately selected parameters reasonably predicted the average daily $NO_2$ concentrations. Including a weight matrix in the objective function resulted in a better data fit.
Our visual inspection of the posterior MCMC samples suggests high levels of agreement and demonstrate
little uncertainty in their predictions.
We also experimented with an LSTM model but could not achieve competitive results.
Although weak for some years, we observe correlations between the SVI and
NO2 concentration, most noticeable in 2010.&lt;/p>
&lt;h2 id="interested-in-learning-more">Interested in Learning More?&lt;/h2>
&lt;p>Please see our &lt;a href="content/2023-NO2-Poster.pdf">poster&lt;/a>.&lt;/p></description></item><item><title>Low-Level Ozone Prediction</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2023-ozone/</link><pubDate>Fri, 01 Dec 2023 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2023-ozone/</guid><description>&lt;!-- https://docs.google.com/document/d/15wNWsSkCo8XK4YvwSOmb3OCASHk8JbV1BKHrsnoi93M/edit -->
&lt;p>This blog post was written by Antonio Gamboa, Lucretia Gant, Timothy Gant, Andrea Glover, and Julianna Long and published with minor edits. The team was advised by &lt;a href="../author/bree-ettinger">Bree Ettinger&lt;/a>.In addition to this post, the team has also given a &lt;a href="content/2023-Ozone-Slides">midterm presentation&lt;/a>, filmed a &lt;a href="https://youtu.be/Cd_eRpne4Ns" target="_blank" rel="noopener">poster blitz video&lt;/a>, created a &lt;a href="content/2023_Ozone-Poster.pdf">poster&lt;/a> and written a &lt;a href="content/2023-Ozone-Manuscript.pdf">manuscript&lt;/a>.&lt;/p>
&lt;h2 id="motivation">Motivation&lt;/h2>
&lt;p>Our team is identifying potential correlations between ground-level ozone pollution and its social and economic disparities, reducing environmental justice. We aim to utilize the results from this investigation to advocate for equitable policies and regulations.&lt;/p>
&lt;p>Bringing social justice to communities impacted by unidentified factors depends on the guidance from scientific research that exposes and discovers inequalities. Communities with lower socioeconomic status often experience a disproportionate burden of air pollution, including higher levels of ground-level ozone, due to factors like proximity to industrial areas, transportation hubs, and poorer air quality monitoring. However, it is unclear how ground-level ozone pollution becomes an issue of social justice and equitable protection of public health for all communities.&lt;/p>
&lt;h2 id="what-is-ozone-and-how-is-it-formed">What is Ozone, and how is it formed?&lt;/h2>
&lt;p>Ozone is a molecule composed of three oxygen atoms; it is highly reactive and mostly found as a colorless and odorless gas. At high altitudes in the stratosphere, it protects us from the damaging UV sun rays&amp;gt; However, at the ground level, in the troposphere, ozone is a secondary pollutant. As a secondary pollutant, ozone, it is formed through complex chemical reactions of molecules in the air.&lt;/p>
&lt;p>Ground-level Ozone (O3) is formed when nitrogen oxides (NOx) and volatile organic compounds (VOCs) react in the presence of sunlight.
NOx is primarily emitted from vehicle exhaust, power plants, and industrial processes, while VOCs come from sources such as gasoline vapors, solvents, and certain industrial activities. These emissions are known as ozone precursors.&lt;/p>
&lt;p>When these precursors mix in the lower atmosphere and are exposed to sunlight, a series of chemical reactions occur, &lt;a href="https://doi.org/10.3389/fimmu.2019.02518" target="_blank" rel="noopener">leading to the formation of ground-level ozone &lt;/a>.&lt;/p>
&lt;h2 id="how-mathematics-may-be-used-to-better-monitor-and-deal-with-ozone">How Mathematics may be used to better monitor and deal with Ozone?&lt;/h2>
&lt;p>Mathematics is currently used to study air pollution. Mathematics provides us with equations and algorithms to solve a problem through computation. The mathematical algorithms act as an exact set of instructions to be followed by the hardware in the computer. The result of such a combination of mathematics and computer science provides us with important functional models that simulate and predict ozone concentrations.&lt;/p>
&lt;p>Mathematical modeling can be done by taking large data sets for ozone collected daily over the years and finding a function that fits the data well with a small percent error. Furthermore, computation techniques such as machine learning can be applied to the data to predict future ozone levels.
The mathematical prediction of low-level ozone can inform air quality forecasts and help people determine which outside activities are safe.&lt;/p>
&lt;h2 id="what-type-of-mathematical-applications-may-be-used-to-better-understand-ground-level-ozone">What type of Mathematical Applications may be used to better understand Ground Level Ozone?&lt;/h2>
&lt;p>Functional linear regression models have recently shown their ability to predict low ozone levels. Specifically, functional models based on &lt;a href="https://onlinelibrary.wiley.com/doi/10.1002/env.2147" target="_blank" rel="noopener">bivariate splines over triangulation&lt;/a> approximate the spatially distributed ozone measurements on a surface. This powerful computation allows the forecast and the quantifying of the uncertainty associated with the predictions of ozone levels for a given day.
A mathematical functional regression model is a statistical technique combining functional data analysis and regression analysis principles. It is used to model the relationship between a dependent variable and one or more independent variables when both are functions or curves rather than simple numeric values. The data points are considered as functions or curves rather than individual observations.
Functional principal component regression combines functional principal component analysis (PCA) and regression analysis. It uses the principal components of the functional independent variables to represent the variation in the data and estimates the coefficients for predicting the functional dependent variable.
In this project, we have used PCA as a powerful tool for analyzing and modeling relationships between ground-level ozone and the SVI index for communities near the Atlanta area.&lt;/p>
&lt;h2 id="how-could-this-modeling-be-used-to-address-social-justice">How could this modeling be used to address Social Justice?&lt;/h2>
&lt;p>Our preliminary results suggest a potential correlation between ground-level ozone
and socially vulnerable neighborhoods. Social Vulnerability Indexes (SVI) identify communities predisposed to harm if exposed to a natural disaster or harmful conditions. These communities are particularly at risk because of specific factors related to their socioeconomic status, household characteristics, racial and ethnic minority status, and housing type. An SVI index is helpful to governments and communities because it identifies and acknowledges which additional support and attention is needed to mitigate the harmful impacts of harmful environmental exposure.
Our model approximates the spatially distributed ozone measurements on a surface, thereby predicting low-level ozone. In addition, by including the SVI and the Environmental Justice Index (EVI), our model calls for a better predictive tool. This tool can inform air quality and help determine which communities are most at risk. Our computational mathematical method also quantifies the uncertainty associated with our predictions and helps us identify the spatial distribution of low-level ozone among the most vulnerable communities.&lt;/p>
&lt;h2 id="interested-in-learning-more">Interested in Learning More?&lt;/h2>
&lt;p>Please see our &lt;a href="content/2023-Ozone-Manuscript.pdf">manuscript&lt;/a> or &lt;a href="content/2023-Ozone-Poster.pdf">poster&lt;/a>.&lt;/p></description></item><item><title>Comparing Reinforcement Learning to Optimal Control Methods on the Continuous Mountain Car Problem</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2022-rl-vs-oc/</link><pubDate>Fri, 05 Aug 2022 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2022-rl-vs-oc/</guid><description>&lt;h2 id="what-is-the-best-way-to-get-a-car-out-of-the-bottom-of-a-hill">What is the best way to get a car out of the bottom of a hill?&lt;/h2>
&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>This blog post was written by &lt;a href="https://www.linkedin.com/in/jacob-mantooth-7b262321b/" target="_blank" rel="noopener">Jacob Mantooth&lt;/a>,&lt;a href="https://dewanchowdhury.github.io/" target="_blank" rel="noopener">Dewan Chowdhury&lt;/a>, &lt;a href="https://www.linkedin.com/in/arjunso/" target="_blank" rel="noopener">Arjun Sethi-Olowin&lt;/a> and published with minor edits.The team was advised by Dr.Lars Ruthotto.In addition to this post, the team has also given a &lt;a href="content/2022_REU_RLvsOC_MidtermPresentation.pdf">midterm presentation&lt;/a>, filmed a &lt;a href="https://www.youtube.com/watch?v=i9g6mRNJEHA&amp;amp;feature=youtu.be" target="_blank" rel="noopener">poster blitz video&lt;/a>, created a &lt;a href="content/2022_REU_RLvsOC_Poster.pdf">poster&lt;/a>, published &lt;a href="https://github.com/EmoryMLIP/MountainCar-RLvsOC" target="_blank" rel="noopener">code&lt;/a>, and written a &lt;a href="">paper&lt;/a>.&lt;/p>
&lt;p>The word around town is that reinforcement learning is the top dog and has the answers to all our problems. We wanted to see if that really was the case, so this summer we took a trip to Emory University where we looked at the continuous mountain car problem to see if reinforcement learning really, was the best. The continuous mountain car problem is an example of an optimal control problem. In the image below you can see an example of what the continuous mountain car problem looks like.&lt;/p>
&lt;p align="center">
&lt;img src="img/mountaincar.png" width="50%" height="50%"/>
&lt;/p>
&lt;p>You may be asking yourself what is an optimal control problem? An optimal control problem is problem that consists of controlling a dynamical system to minimize (or maximize) a given objective function. In our case the continuous mountain can be modeled as a ODE where $u$ is some controllable function. In the continuous mountain car problem our control,$u$, is whether the car accelerates left or right. In an optimal control problem, we seek to optimize some objective function, in our case we will minimize the objective function. Our two objective functions are the running cost, which penalizes the car for acceleration. While our other objective function is the terminal cost which penalizes the car for not reaching the goal, the top of the hill, in time.&lt;/p>
&lt;h3 id="why-this-problem">Why This Problem?&lt;/h3>
&lt;p>You may be wondering why choose the Continuous Mountain Car Problem? Here are a couple of reasons why we picked this example,&lt;/p>
&lt;ul>
&lt;li>
&lt;p>Established benchmark for RL models&lt;/p>
&lt;/li>
&lt;li>
&lt;p>2-D state-space allows for good plots and visualizations&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Both RL and optimal control problem&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Finite horizon (time)&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Continuous state and motion&lt;/p>
&lt;/li>
&lt;/ul>
&lt;p>The whole reason we are doing this is because we want to compare three different ways of solving the continuous mountain car problem and see which one really is the best. Our three approaches are&lt;/p>
&lt;ul>
&lt;li>
&lt;p>Local solution using numerical ODE solvers and nonlinear optimization (baseline)&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Reinforcement learning with actor-critic algorithm (data-based approach)&lt;/p>
&lt;/li>
&lt;li>
&lt;p>Optimal control using both model and data&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h2 id="our-three-approaches">Our Three Approaches&lt;/h2>
&lt;h3 id="a-local-method">A Local Method&lt;/h3>
&lt;p>Our first method that we looked at during this REU was the local method. We tried to find the optimal control $u_h$ by formulating an optimization problem.&lt;/p>
&lt;p>We first discretize the control, state and the Lagrangian.&lt;/p>
&lt;p>Setting $z_h^{(0)}=z_t$ and $\ell_h^{(0)}=0$, allows us to use a forward Euler scheme for some control $u$.&lt;/p>
&lt;p>We then approximate our objective function which yields the optimization problem&lt;/p>
&lt;p>To solve our optimization problem, we used gradient descent. By taking an initial guess for $u_h$ and repeatedly updating $u_h$ using the gradient of the objective function and step size $\alpha$&lt;/p>
&lt;div style="text-align: center">
&lt;p>$(u_h)_0 = \vec0$&lt;/p>
&lt;p>$\vdots$&lt;/p>
&lt;p>$(u_h)_6 = (u_h)_5 - \alpha( \nabla J((u_h)_5))$&lt;/p>
&lt;p>$\vdots$&lt;/p>
&lt;p>$(u_h)_* = (u_h)_19 - \alpha( \nabla J((u_h)_19))$&lt;/p>
&lt;/div>
&lt;p>Below is a nice visual example of what all this math means. When the tail reaches the dotted line, it means our car has reached the top of the hill.&lt;/p>
&lt;p align="center">
&lt;img width="460" height="300" src=img/pvsv_color_local.gif>
&lt;/p>
&lt;p>The graph shows us the position vs velocity of the car. In the graph the black dot represents t and the tail of the plot, when x-position is .45 is time T. In the graph we see the color change from red-blue, in our plot the blue color is when the car is accelerating, and control is positive but red otherwise.&lt;/p>
&lt;p>Our next goal was how do we create a nice visualization of what the actual solution looks like.&lt;/p>
&lt;p align="center">
&lt;img src="img/Localmethod.gif" >
&lt;/p>
&lt;p>We see in this video what our optimal local solution looks like. A couple of things that should be noted is, if we move the car to a new position then this local solution may no longer work. The same can be said if we slowed down/speed up a bit then this solution may not even let the car get to the top of the mountain. Another downside of the local is that it is a non-linear and non-convex problem which makes this method slow. This local solution will serve as a baseline so we can compare other methods to something to see which one is really the best.&lt;/p>
&lt;h2 id="global-methods">Global Methods&lt;/h2>
&lt;p>Now that we have established a baseline, we will discuss our other two methods. Our other two methods that we will be looking at are global methods, the first being reinforcement learning method and the other being optimal control method. You may be asking yourself what is the difference? Reinforcement learning is more of a data driven approach while the optimal control method is a hybrid approach, using both a model and data.&lt;/p>
&lt;h1 id="reinforcement-learning">Reinforcement Learning&lt;/h1>
&lt;p>Our first stop in exploring global methods is reinforcement learning. We will be using reinforcement learning with actor-critic algorithm. This approach is completely data-based approach. In reinforcement learning it has no knowledge of the model, it only considers the objective function. In reinforcement learning we would like to maximize a reward, so in our case we will maximize negative cost. Reinforcement learning is stochastic in two ways with initial position and action space which allows for exploration. In Reinforcement learning we are trying to estimate an optimal control policy. One of the big things that we have yet to discuss is, what is actor-critic algorithm ? The actor-critic (AC) architecture for RL is well-suited for a continuous action-space as in the continuous mountain car problem &lt;strong>&lt;a href="https://arxiv.org/abs/1509.02971" target="_blank" rel="noopener">1&lt;/a>&lt;/strong> . In the actor-critic algorithm the critic must learn about and critique whatever policy is currently being followed by the actor. We worked in the OpenAI gym mountain car environment, so we were able to find preexisting code for our project. We also were able to adapted the TD advantage actor-critic algorithm adapted from &lt;strong>&lt;a href="https://medium.com/@asteinbach/actor-critic-using-deep-rl-continuous-mountain-car-in-tensorflow-4c1fb2110f7c" target="_blank" rel="noopener">here&lt;/a>&lt;/strong>.
The thing is our preexisting code was not the same as our problem, so we had to modify it some. After some modification to the code, the following video is the results that we were able to get after many training cycles.&lt;/p>
&lt;p align="center">
&lt;img src="img/RLmethod.gif" >
&lt;/p>
&lt;p>In this video we see that RL gave us a sub optimal solution compared to the local solution. You may also notice that in the reinforcement learning method our car takes an extra swing backwards to get to the top of the hill. In our reinforcement learning method it took 1000&amp;rsquo;s episodes just to get the car to our goal. We saw that reinforcement learning is very fragile, a couple of changes saw our success rate go from 70% to barely making it all. The picture below is position vs velocity of the car.&lt;/p>
&lt;p align="center">
&lt;img src="img/pvsv_color_rl-1.png" width="50%" height="50%"/ >
&lt;/p>
&lt;p>as you can see compared to the local solution, we see that the RL solution is very sub optimal solution.&lt;/p>
&lt;h1 id="optimal-control-method">Optimal control method&lt;/h1>
&lt;p>Our last two methods were vastly different with reinforcement learning using a data driven approach and the local method using a model-based approach. We will now be looking at the optimal control method which combines both model and data driven approaches. In this approach, we aim to adhere to the method discussed &lt;strong>&lt;a href="https://arxiv.org/abs/2104.03270" target="_blank" rel="noopener">here&lt;/a>&lt;/strong>. We will estimate the corresponding value function with neural network approximators utilizing feedback from the Hamilton-Jacobi-Bellman equation and Hamiltonian.&lt;/p>
&lt;p>Using OC method we were able to produce the following&lt;/p>
&lt;p align="center">
&lt;img src="img/pvsv_color_oc-1.png" width="50%" height="50%"/ >
&lt;/p>
&lt;p>Once again, we created a position vs velocity of the car graph. As you can see this graph is a sub optimal solution compared to the local method. Compared to the RL method, we see how much better OC was for our problem. We see through testing of the RL method that there are some draw backs to forgetting the model and just being purely driven by data.&lt;/p>
&lt;h2 id="our-experiences">Our Experiences&lt;/h2>
&lt;h3 id="week-1">Week 1&lt;/h3>
&lt;p>In week one we decided to make a game plan for the following weeks. We would work on the local method for just a week since it was basically finished. For the other two methods we would spend two weeks on each method. Lastly, we would save the last week to wrap up all three methods and anything else that is left over. During the first week we wanted to look at the local method and explore it some.&lt;/p>
&lt;h3 id="week-2--3">Week 2 &amp;amp; 3&lt;/h3>
&lt;p>We decided to spend two weeks to look at reinforcement learning, during these two week we were able to produce a PowerPoint in beamer for our mid-week presentation. In week two and three we looked at our first global method. Dr.Ruthotto handed us some pre written code to play around with. A thing that should be noted is that this prewritten code only worked maybe 50% of the time. Looking deeper into the code we realized that we would have to mess around with the code to get it to match our problem. After we made these couple of changes in our code, we saw how fragile reinforcement learning is, instead of working 50% of the time our code barely worked at all. During this week we were also able to produce a rendering of both of our local and RL methods, adding a nice touch to our presentation that we gave.&lt;/p>
&lt;h3 id="week-4--5">Week 4 &amp;amp; 5&lt;/h3>
&lt;p>During week 4 and 5 we looked at optimal control method. During week 4 we were able to produce a rough draft of the website while also taking a deeper look into optimal control method. Week five we created a rough and final draft of our poster. While working on our poster we were able to produce a graph for OC so we can compare it to our other methods.&lt;/p>
&lt;h3 id="week-6">Week 6&lt;/h3>
&lt;p>We wrapped up any unfinished work including our paper, OC method and this website. After struggling with code for method three we finally found that our OC method had better results than RL. During week six we also gave a poster talk to the Emory staff and students. Our group ended up winning best poster and we each won some amazon gift cards. We will continue working on method three to try and get it to work for 150 steps, we also look to fine tune our code for method three and two.&lt;/p>
&lt;h2 id="more-about-our-team">More About Our Team&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>&lt;a href="https://www.linkedin.com/in/jacob-mantooth-7b262321b/" target="_blank" rel="noopener">Jacob Mantooth&lt;/a>&lt;/strong>&lt;/li>
&lt;li>&lt;strong>&lt;a href="https://dewanchowdhury.github.io/" target="_blank" rel="noopener">Dewan Chowdhury&lt;/a>&lt;/strong>&lt;/li>
&lt;li>&lt;strong>&lt;a href="https://www.linkedin.com/in/arjunso/" target="_blank" rel="noopener">Arjun Sethi-Olowin&lt;/a>&lt;/strong>&lt;/li>
&lt;/ul>
&lt;h2 id="reference">Reference&lt;/h2>
&lt;p>U. M. Ascher and C. Greif. &amp;ldquo;Chapter 9 (Optimization).&amp;rdquo; A First Course on Numerical Methods. SIAM. SIAM,2011&lt;/p>
&lt;p>U. M. Ascher and C. Greif. &amp;ldquo;Chapter 14 (Numerical time integrators).&amp;rdquo; A First Course on Numerical Methods. SIAM. SIAM,2011&lt;/p></description></item><item><title>Undergraduate students and teacher develop segmentation methods for diagnosing Chiari Malformation</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2021-chiari/</link><pubDate>Tue, 14 Dec 2021 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2021-chiari/</guid><description>&lt;p>This post was written by Justin Smith, Elle Buser, Emma Hart, and Ben Hueneman and published with minor edits. The team was advised by Dr. Lars Ruthotto.
In addition to this post, the team has also created slides for a &lt;a href="https://github.com/EmoryMLIP/emory-reu-ret-website/blob/main/content/projects/2021-chiari/Chiari_Disease_Presentation.pdf" target="_blank" rel="noopener">midterm presentation&lt;/a>, a &lt;a href="https://youtu.be/tdjXj3JdpQU" target="_blank" rel="noopener">poster blitz&lt;/a> video, &lt;a href="https://github.com/lruthotto/ChiariProject" target="_blank" rel="noopener">code&lt;/a>, and a &lt;a href="https://arxiv.org/abs/2109.14116" target="_blank" rel="noopener">paper&lt;/a>.&lt;/p>
&lt;h2 id="collaboration-never-sleeps">Collaboration Never Sleeps.&lt;/h2>
&lt;p>During our summer research at Emory University 2021 REU/RET program, our group focused on the algorithmic diagnosis of Chiari malformation from DENSE MRIs. We created an algorithm that can accurately and efficiently segment the cerebellum and brain stem from a magnitude image and use displacement data to classify whether or not a patient has the Chiari malformation. In doing so, we investigated two approaches; one that segments the given image by aligning and comparing the image to a known atlas and another that segments through deep learning.&lt;/p>
&lt;h2 id="did-somebody-say-chiari-malformation">Did Somebody Say Chiari Malformation?&lt;/h2>
&lt;p>Chiari malformation is a condition in which brain tissue extends into the spinal canal. While it can be difficult to diagnose Chari from anatomical images, a promising new direction for diagnosis is by looking at brain movement . Using an MRI technique called DENSE (shown below) that records how the brain moves, &lt;a href="https://link.springer.com/article/10.1007/s10439-020-02695-7" target="_blank" rel="noopener">Dr. Oshinsky’s group (at Emory&amp;rsquo;s Dept. of Radiology)&lt;/a> collected data about how Chiari patients have more brain movement in the cerebellum and brainstem than controls.&lt;/p>
&lt;img src="img/DENSE.gif" alt="DENSEgif" width="500"/>
&lt;p>This method may be more accurate in diagnosing Chiari, however, the large number of manual processing steps may limit its use as a wide-spread screening tool. This project aims at exploring the use of machine learning algorithms to automize parts of the image processing pipeline, most critically the segmentation of the image into different brain regions. The teams worked with image data that has been collected and labeled by Dr. Oshinski’s group in a previous research study. The project is accessible to the team members since we can build upon recent progress and software made in image processing and computer vision and the image data is two-dimensional and of limited resolution, which enables fast experimentation. Despite this simplicity, the project allows us to investigate ML in a realistic setting and investigate the generalization properties and robustness of the approach.&lt;/p>
&lt;h2 id="leave-the-segmentation-to-us">Leave the SEGMENTATION to US!&lt;/h2>
&lt;img src="img/five-masks.png" alt="img/Chiari-Synergy" width="800"/>
We develop this project to solve the problem of identifying where the brain stem and cerebellum are in a given MRI. By finding or, in the language of the field, by segmenting the brain stem and cerebellum, we find the most relevant regions to look at brain movement. Using the DENSE MRI data, we can then average the movement over those regions to produce a biomarker that can help predict whether or not a patient has the Chiari Malformation. By producing these segmentations (examples above) automatically with the machine learning or atlas-based approaches, the diagnosis process could become much cheaper and more efficient.
&lt;h2 id="atlas-based-image-registration-vs-machine-learning">Atlas Based Image Registration vs Machine Learning.&lt;/h2>
&lt;p>We first looked into atlas-based image registration as a way to produce automatic segmentations of the brain stem and cerebellum. Using the FAIR toolbox in MATLAB, the idea behind this method was to have a bank of MRI images with manually drawn segmentations that we could compare a new MR image to. Once we find a transformation (example below) between the known and new images, we can use the same transformation to produce a new segmentation from the know one.&lt;/p>
&lt;img src="img/AtlasGIF.gif" alt="registration" width="500"/>
&lt;p>We also looked into a machine learning approach. The goal here was to find a relationship between the DENSE images and their corresponding manual segmentations by training a model using convolutional neural networks (CNN). The network &amp;quot;learns&amp;quot; to identify images features, and, if successful, DENSE images can be used as inputs and the model will automatically segment the brain stem and cerebellum.&lt;/p>
&lt;img src="img/MachLearningDiagram.jpg" alt="MachLearningDiagram" width="500"/>
&lt;p>Our project implemented a CNN called U-Net, you can find out more about this network and the code we used here: &lt;a href="https://amaarora.github.io/2020/09/13/unet.html" target="_blank" rel="noopener">U-Net: A PyTorch Implementation in 60 lines of Code&lt;/a>&lt;/p>
&lt;p>Overall, we found that the machine learning method produces better results, both in terms of segmentations of the brain stem and cerebellum, and in terms of how accurate the biomarkers calculated from those segmentations are. Results were very close to the manually produced target results, and we have ideas for further work that could make them even closer!&lt;/p>
&lt;p>Additional information: to learn about atlas-based image registration and machine learning, check out these links!
&lt;a href="https://www.sicara.ai/blog/2019-07-16-image-registration-deep-learning" target="_blank" rel="noopener">What is Image Registration?&lt;/a>
&lt;a href="https://youtu.be/QghjaS0WQQU" target="_blank" rel="noopener">What is Machine Learning?&lt;/a>&lt;/p>
&lt;h2 id="time-management-is-everything">Time Management is Everything!&lt;/h2>
&lt;p>Here&amp;rsquo;s an outline of our process, as it evolved with time.&lt;/p>
&lt;p>&lt;strong>Week 1:&lt;/strong>
During our first week, we created a working atlas-based image registration example, using FAIR: a MATLAB image registration toolset. We also began to look at the at a machine learning method, called U-Net, that we began setting up using PyTorch. We also set up GitHub, that we used throughout the project to collaborate on and publish codes.&lt;/p>
&lt;p>&lt;strong>Week 2:&lt;/strong>
One of the first things we noticed when we began working with our data set the first week was the variability in the brightness and contrast of our images. In week two, we explored some different methods to help enhance the images. We decided to use a tool in MATLAB&amp;rsquo;s Image Processing Toolbox to normalize the images in a process called &lt;a href="http://www.sci.utah.edu/~acoste/uou/Image/project1/Arthur_COSTE_Project_1_report.html" target="_blank" rel="noopener">histogram normalization&lt;/a>, which made our images much more consistent.&lt;/p>
&lt;img src="img/compare.png" alt="normalization" width="500"/>
&lt;p>With this process complete, we began working on other atlas-based examples, and setting up the neural network with default parameters.&lt;/p>
&lt;p>&lt;strong>Week 3:&lt;/strong>
We spent most of the third week preparing for our midterm presentation. It was helpful to practice presentation skills and get familiar using &lt;a href="https://www.overleaf.com/learn/latex/Beamer" target="_blank" rel="noopener">Beamer&lt;/a> - a math presentation tool we weren&amp;rsquo;t yet familiar with that is the gold standard for mathematic presentations - and also to reflect on the progress we made, and talk through next steps with others after we presented. After our presentation, we started refining the machine learning method to use a new optimizer that automatically chooses the algorithm&amp;rsquo;s &lt;a href="https://machinelearningmastery.com/understand-the-dynamics-of-learning-rate-on-deep-learning-neural-networks/" target="_blank" rel="noopener">learning rate&lt;/a> using a &lt;a href="https://machinelearningmastery.com/line-search-optimization-with-python/#:~:text=The%20line%20search%20is%20an,with%20one%20or%20more%20variables.&amp;amp;text=Linear%20search%20is%20an%20optimization%20algorithm%20for%20univariate%20and%20multivariate%20optimization%20problems" target="_blank" rel="noopener">line search&lt;/a>. This produced much better results!&lt;/p>
&lt;img src="img/optimCompare.png" alt="optimCompare" width="500"/>
&lt;p>&lt;strong>Week 4:&lt;/strong>
In the fourth week, we began working on our paper manuscript. Although there was still much of the research left to be done, coding in MATLAB and PyTorch, starting on this process of outlining our problem and explaining our methods helped us clarify next steps. At the end of the week, we turned back to our codes. On the atlas-based side, we explored ways to choose the images to compare to each other, because it can make a big difference in how good the outputted segmentation is. We also began devising a way to compare a new image we want to segment to more than one images we already have segmentations before, hoping that by averaging those results together, we could produce even better segmentations. On the machine learning side, we looked at ways to figure out the ideal number of iterations to use, so that we were neither overfitting nor underfitting our model to training data.&lt;/p>
&lt;p>&lt;strong>Week 5:&lt;/strong>
Week five progressed very similarly to week four, as we continued working on our paper and developing better methods. We also began working on a poster to present our final results more broadly - and on this website!&lt;/p>
&lt;p>&lt;strong>Week 6:&lt;/strong>
In the last week of the program, we had a lot left to do, finishing up our poster, website, and paper manuscript. We finalized every parameter in each method before finally looking at the testing set for the first time and analyzing our final results.&lt;/p>
&lt;img src="img/too-cute-pic.png" alt="brain-heart" width="100"/>
&lt;h2 id="poster-blitz-video">Poster Blitz Video&lt;/h2>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/tdjXj3JdpQU" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen>&lt;/iframe>
&lt;h2 id="more-about-the-team">More About the Team&lt;/h2>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Elle Buser&lt;/strong> is a rising senior at the University of Wyoming majoring in mathematics and minoring in physics. At UWyo, she has worked as a tutor/supplemental instructor for Calc II and as a research assistant in the physics and astronomy department. Outside of school, she spends her time reading, doing arts &amp;amp; crafts, hiking, running, or just enjoying the great outdoors.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Emma Hart&lt;/strong> is a rising senior at Colgate University, double majoring in applied mathematics and educational studies. She enjoys combining these interests in math and education by volunteering as a math tutor and also by working for Colgate as a Writing Center peer consultant and computational mathematics grader. Outside of school, she enjoys crafting and spending time outside with friends and family.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Ben Huenemann&lt;/strong> is a rising junior at the University of Utah majoring in mathematics and computer science. He loves all things math, but is mainly interested in pure mathematics and plans to go on studying algebra or something similar in graduate school. Outside of school, he enjoys playing piano, evening bike rides, and board games with friends/family.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Justin Smith&lt;/strong> is an Early Childhood educator in Atlanta, Georgia, where he is in his 10th year of teaching. He is certified in Early Childhood Education and holds a Gifted Endorsement in the state of Georgia. He currently holds the Teacher of the Year award at his school. He is very active in his community and school, where started an organization called Obama Gentlemen with the primary duty to improve academics and address behavioral concerns of young men by instilling morals and values that will help transition young adolescent males into young adulthood. During his leisure time, he enjoys traveling, shopping, listening to music, and spending time with his family.&lt;/p>
&lt;/li>
&lt;/ul></description></item><item><title>AI-Assisted Exploration of the Polygonal Faber-Krahn Inequality</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2026-faber-krahn/</link><pubDate>Fri, 16 Jan 2026 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2026-faber-krahn/</guid><description>&lt;p>&lt;strong>Mentor:&lt;/strong> Dr. Levon Nurbekyan&lt;/p>
&lt;h2 id="overview">Overview&lt;/h2>
&lt;p>This project focuses on the &lt;em>polygonal Faber-Krahn inequality&lt;/em>, which conjectures that among all $n$-gons of fixed area, the regular $n$-gon minimizes the first Dirichlet eigenvalue of the Laplacian. The conjecture was introduced by Polya and Szego, who proved it for $n=3,4$ using Steiner symmetrization. For $n \ge 5$, however, these techniques break down, and the conjecture remains open.&lt;/p>
&lt;p>The goal of the project is to use AI-assisted methods to explore this problem both computationally and conceptually, with the aim of either identifying potential counterexamples or producing empirically supported conjectures and strategies that could inform future theoretical work.&lt;/p>
&lt;h2 id="computational-approach">Computational Approach&lt;/h2>
&lt;p>For a broad range of values of $n$, we will numerically search for candidate optimal $n$-gons using a combination of gradient-based optimization and reinforcement learning, and subsequently leverage evolutionary coding agents in the spirit of AlphaEvolve to refine and extend these optimization strategies. The resulting large-scale experiments will either provide numerical evidence supporting the conjecture or identify configurations that warrant closer scrutiny.&lt;/p>
&lt;h2 id="analysis">Analysis&lt;/h2>
&lt;p>In the absence of apparent counterexamples, we will use models with advanced reasoning capabilities, in the spirit of DeepThink, to analyze the numerical results at a higher level, seeking recurring patterns or structural features in the evolution of optimizing shapes that may suggest new geometric or analytical approaches to the problem.&lt;/p>
&lt;h2 id="prerequisites">Prerequisites&lt;/h2>
&lt;p>A solid background in linear algebra, multivariable calculus, and basic numerical methods.&lt;/p></description></item><item><title>New algorithm helps make portable CT machine images clearer</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2021-tomography/</link><pubDate>Tue, 14 Dec 2021 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2021-tomography/</guid><description>&lt;p>This blog post was written by Mai Phuong Pham Huynh, Manuel Santana, and Ana Castillo and published with minor edits. The team was advised by Dr. James Nagy.
In addition to this post, the team has also given a &lt;a href="https://github.com/EmoryMLIP/emory-reu-ret-website/files/6874766/_REU2021__Tomo_Presentation.pdf" target="_blank" rel="noopener">midterm presentation&lt;/a> , created a &lt;a href="./img/REU_RETPoster.pptx.png">poster&lt;/a>, made a &lt;a href="https://youtu.be/qdcGe9MKCoI" target="_blank" rel="noopener">poster blitz video&lt;/a>, published &lt;a href="https://github.com/manuelarturosantana/TomoREU2021" target="_blank" rel="noopener">code&lt;/a>, written a &lt;a href="https://arxiv.org/abs/2109.01481" target="_blank" rel="noopener">paper&lt;/a>, and developed a &lt;a href="Tomography.Lesson.pdf">lesson plan&lt;/a>.&lt;/p>
&lt;p>Research done at the 2021 REU/RET summer program at Emory University has devised a way to make medical diagnoses more effectively. The work behind this new algorithm relies heavily upon mathematical techniques that connect numerical linear algebra with optimization methods.&lt;/p>
&lt;p>Some of the most important aspects of this work link branches of mathematics like calculus, statistics, algebra and geometry. The different contributions by mathematicians to these branches of mathematics were complementary and accelerated research in the area of image reconstruction. The results obtained from their work have offered other student mathematicians a different way to study these problems and shown that this is an effective way for radiology doctors to accurately diagnose illnesses and keep saving patients&amp;rsquo; lives. The goal of the Point-of-Care Tomographic Imaging research group is to develop numerical methods that would estimate the geometry parameters of a portable CT scanning device and to reconstruct the image. CT images in portable devices are a lot clearer. This work has been one of the most interesting pieces of work ever accomplished.&lt;/p>
&lt;h1 id="covid-19-and-imaging">COVID-19 and Imaging&lt;/h1>
&lt;p>The mathematical ideas involved in the Point-of-Care Tomographic Imaging project are of great importance. With these mathematical methods, imaging becomes a possibility with patients in all parts of the world, allowing them to live a healthy and happy life. To get better reconstructed images, the group considers a regularized linear least squares problem, to which the solution is approximated by using an alternating descent scheme known as block coordinate descent or (BCD). Each iteration of BCD has two steps. One BCD step solves a linear least square problem, and the next step solves the nonlinear problem.&lt;/p>
&lt;p style="font-size: 0.9rem;font-style: italic;"> &lt;img style="display: block;" src="https://live.staticflickr.com/65535/49679288857_200a679135_b.jpg" alt="Coronavirus">&lt;a href="https://www.flickr.com/photos/110751683@N02/49679288857">"Coronavirus"&lt;/a>&lt;span> by &lt;a href="https://www.flickr.com/photos/110751683@N02">Yu. Samoilov&lt;/a>&lt;/span> is licensed under &lt;a href="https://creativecommons.org/licenses/by/2.0/?ref=ccsearch&amp;atype=html" style="margin-right: 5px;">CC BY 2.0&lt;/a>&lt;a href="https://creativecommons.org/licenses/by/2.0/?ref=ccsearch&amp;atype=html" target="_blank" rel="noopener noreferrer" style="display: inline-block;white-space: none;margin-top: 2px;margin-left: 3px;height: 22px !important;">&lt;img style="height: inherit;margin-right: 3px;display: inline-block;" src="https://search.creativecommons.org/static/img/cc_icon.svg?image_id=6f5e7d78-bf4e-4b46-8ba9-c469af85860f" />&lt;img style="height: inherit;margin-right: 3px;display: inline-block;" src="https://search.creativecommons.org/static/img/cc-by_icon.svg" />&lt;/a>&lt;/p>
&lt;p>By the 1900’s mathematicians had already found ways to use these iterative reconstruction techniques to reconstitute an image. In 1917, Johann Radon introduced the Radon Transform, followed by Stefan Kaczmarz in 1937. Scientists Allan McLeod Cormack and Godfrey Newbold Hounsfield developed the first scanning device. These innovations each contributed to modern medical scanning technology.&lt;/p>
&lt;p>The medical field has greatly depended on these iterative reconstruction methods. In medical imaging, computed tomography (CT) techniques are becoming more and more popular for their ability to produce high quality images of the human body. CT methods use a combination of computer processes and mathematics to reconstruct images. The COVID-19 pandemic brought many challenges and setbacks to many doctors around the world and imaging played a vital role in helping diagnose the virus. Doctors have been able to use CT scans to diagnose and treat COVID-19 by examining the lungs of patients who are potentially infectious, or in recovery.&lt;/p>
&lt;h1 id="point-of-care-tomographic-imaging">Point-of-Care Tomographic Imaging&lt;/h1>
&lt;p>Since the first CT scanner was developed, CT scanning technology has evolved significantly. Now, there are portable CT scanners that can be transported anywhere without trouble. The idea of transporting huge heavy CT scanners to remote locations was always a difficult task for many doctors. Point-of-care tomographic imaging has allowed radiologists to add portable CT scanners to their departments to increase patient satisfaction and improve medical outcomes. Portable CT scanners can be used to do scans of a patient without moving the patient out of bed.&lt;/p>
&lt;p>A CT scanner is a device that is composed of a scanning gantry, x-ray generator, computer system, console panel and a physician’s viewing console. The scanning gantry is the part that will produce and detect x-rays. In a typical CT scan, a patient lays on a bed that moves through the gantry. An X-ray source rotates around the patient and shoots X-ray beams through the human body at different angles. These X-ray measurements are then processed on a computer using mathematical algorithms to create tomographic (cross-sectional) images of the tissues inside the body.&lt;/p>
&lt;p>Limitations arise when using these portable CT scanners for medical procedures since these devices require extensive care, such as regular calibration for effective performance. This is where the point-of-care tomographic imaging problem begins.&lt;/p>
&lt;h1 id="the-beginning-of-the-problem">The Beginning of the Problem&lt;/h1>
&lt;p align="center">
&lt;img width="300" height="300" src="https://user-images.githubusercontent.com/84742324/126915305-c8b40ba1-d37f-4317-9bf6-66e5155a8cd5.png">
&lt;/p>
&lt;p>Two parameters that relate to the geometry of portable CT scanning devices are:&lt;/p>
&lt;ol>
&lt;li>R - distance between source and detector and&lt;/li>
&lt;li>θ - orientation of source to detector.&lt;/li>
&lt;/ol>
&lt;p>During the CT acquisition process, the x-ray source rotates around the patient. Each time the x-ray source moves to a new position, these parameters change. In point-of-care imaging, these parameters are essential during the data imaging process. Imprecise data (unknown parameters) result in the reconstruction of poor quality images, while precise data (known parameters) yield a well reconstructed image.&lt;/p>
&lt;p>&lt;em>How were these geometry parameters estimated to obtain better reconstructed images?&lt;/em>&lt;/p>
&lt;p align="center">
&lt;img width="700" height="300" src="https://user-images.githubusercontent.com/84742324/126884434-a47ee15e-2df5-45f1-8b5f-f1af63e894ad.png">
&lt;/p>
&lt;p>The group considered solving a regularized linear least squares problem. The solution to this problem is approximated by using an alternating descent scheme known as block coordinate descent or (BCD). They used this alternating algorithm in which each iteration alternately solves a linear last squares problem (for the image) and a nonlinear least squares problem (for the geometry parameters). An alternate form of this algorithm was also studied as a fixed point iteration, with acceleration techniques used to improve convergence. Implementation and numerical examples were done using MATLAB to obtain results. To understand the background behind this problem, the connection between Beer’s Law and linear algebra needs to be studied.&lt;/p>
&lt;h1 id="beers-law--linear-algebra">Beer’s Law &amp;amp; Linear Algebra&lt;/h1>
&lt;p>The Beer-Lambert Law or Beer’s Law was first developed in 1729 by Pierre Bouguer, and in 1852, Beer added to the law. Beer’s law states that as an x-ray source emits an x-ray beam and it passes through an object, the beam loses energy exponentially. To create a 2-D image, an imaginary grid of pixels is placed atop the object, and an x-ray beam is then shot through the object. Beer’s Law is used to create a system of linear equations.&lt;/p>
&lt;p align="center">
&lt;img width="400" height="300" src="https://user-images.githubusercontent.com/84742324/126884449-423a5fc2-3aa9-4d37-84ca-a8d2c73e3f5a.png">
&lt;/p>
&lt;p>From Algebra, it is known that a system of equations has three types of solutions, infinitely many, exactly one, or no solution. When there is no solution, the system is called inconsistent, as opposed to consistent when there exists a solution. When there are more equations (x-rays) than unknowns (pixels), then the system is overdetermined. This is what happens in computed tomography. When there are more unknowns (pixels) than equations (x-rays), then the system is underdetermined.
In each equation, every unknown represents a physical property, called attenuation, for each pixel value. Pixels that have the same amount of material have the same attenuation. An image is then created from the pixels. The images obtained based on the attenuation values allow doctors to make more accurate diagnoses. For example, bones have a different attenuation than lungs.&lt;/p>
&lt;p>A system of linear equations can be solved in Precalculus by the Gauss-Jordan elimination method. In linear algebra, a matrix equation of the form Ax=b is solved to find x. Ideally, this can easily be done. Unfortunately, for this project, this might not be possible for several reasons: 1) b has noise in it, and 2) the exact matrix A may not be known because the parameters R and θ may be perturbed (for example, the CT machine may not be calibrated correctly). Matrix A may also be a huge matrix containing millions of rows and columns. Typical approaches learned in Pre-Calculus or a linear algebra class cannot be used to solve this problem. The tomographic imaging research group uses other methods to solve these large problems, referred to as ill-posed.&lt;/p>
&lt;p>An ill-posed problem is often referred to as one that is not well-posed. In the 20th Century, French mathematician Jacques Hadamard defined a well-posed problem as one having three properties: 1) a solution exists, 2) the solution is unique, 3) the solution depends continously on initial values. An ill-posed problem needs to be approached differently, e.g. -regularization. Tikhonov regularization is a common form of regularization for ill-posed problems.&lt;/p>
&lt;h1 id="the-linear--nonlinear-least-squares-problem">The Linear &amp;amp; NonLinear Least Squares Problem&lt;/h1>
&lt;p>Throughout the 1800s, mathematicians such as Legendre, Gauss, and Laplace contributed ideas to the method of least squares. For example, the method of least squares can be used to find the line that minimizes the sum of squared distances between teh line and the given data points.
In computed tomography, a problem arises that involves least squares, a regularized linear least squares problem. This regularized linear least squares problem is composed of x, the image, A(p), a matrix that is constructed by the geometry parameters, R and θ, b, the measured data, and a parameter α, called the regularization parameter. The nonlinear least squares problem does not include the regularization parameter.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="./BCDgif.gif" alt="Example of BCD" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>The solution to this problem is found by the block coordinate descent (BCD). The BCD algorithm is an optimization algorithm that solves this problem iteratively. BCD works by minimizing the parameters R and θ and x one at a time. The linear least squares problem is considered first. It takes the initial parameters R and θ, generate matrix A, to get x, the image. Here the parameters, R and θ are known and x is approximated. Then, the nonlinear least squares problem is considered. Once x is known, then x is used to approximate the parameters, R and θ. In this latter phase, x is known and the parameters R and θ are approximated. The pattern continues until a good image is produced.&lt;/p>
&lt;h1 id="filter---based-regularization">Filter - Based Regularization&lt;/h1>
&lt;p>Since it is impossible to get data exactly from the detector due to numerous reasons, regularization is needed to lower the error the solution caused by noise in the data. Tikhonov regularization is one of the most common ways to handle these noise levels. The method is named after mathematician Andrey Tikhonov. Tikhonov worked in numerous topics and different fields of mathematics. His best contributions are in topology.&lt;/p>
&lt;p>Singular value decomposition (SVD) plays an important role in understanding the basic idea behind Tikhonov Regularization. The singular values of the matrix A are significant. To achieve good reconstructed images with small error, it is necessary to have the balance of large and small singular values; the large singular values have good information about the solution, but the small singular values cause errors to blow up.&lt;/p>
&lt;h1 id="matlab-toolboxes--implementation">MATLAB Toolboxes &amp;amp; Implementation&lt;/h1>
&lt;p>MATLAB was used in the project, including the optimization, signal processing, and image processing toolboxes. The IR Tools, AIR Tools II, and Imfil packages were also used.&lt;/p>
&lt;p>To simulate a problem, the IR tools package is used. IR tools also provide regularized linear least square solvers. Three of these linear least squares solvers are: 1) hybrid-LSQR algorithm, 2) IRN, 3) FISTA. To approximate the solution of the non-linear least squares problem two methods are used: 1)lsqnonlin, 2) imfil.&lt;/p>
&lt;p>The implementation of BCD used the IR Tools package and naming conventions. In the base IRtools package, the function PRset is used to set up options for computed tomography problems, where PRset is updated to accept values related to the BCD (inital guess for parameters). To simulate a CT problem with unknown geometry parameters PRtomo_var is used with the image size being n×n. PRtomo_var generates all the data necessary to simulate the inverse problem. IRset updates to include parameters for BCD, such as which acceleration technique to use, and IRbcd implements the BCD algorithm.&lt;/p>
&lt;p>The steps above permit the simulation and solving of a CT reconstruction experiment. With n=64 and perturbations added to the parameters R and θ, the test problem is generated. The image used is the Shepp-Logan image.&lt;/p>
&lt;p align="center">
&lt;img width="800" height="300" src="https://user-images.githubusercontent.com/84742324/126915614-ddf37d04-2973-4b8a-bb23-e086522b99bc.png">
&lt;/p>
&lt;p>After solving the problem with the BCD algorithm, it can be seen that the solution with the BCD resembles that of the true parameters solution.&lt;/p>
&lt;p align="center">
&lt;img width="600" height="300" src="https://user-images.githubusercontent.com/84742324/126915629-d2eb8ef4-77d6-4375-a404-6e3d1876e1bf.png">
&lt;/p>
&lt;h1 id="acceleration-techniques-tests--results">Acceleration Techniques: Tests &amp;amp; Results&lt;/h1>
&lt;p>Other mathematicians continue their work in these fields. Some of them introduced ideas on acceleration techniques to speed up the convergence rate of fixed point iteration problems. One of them is Donald G. Anderson, famous for his work on the Anderson Acceleration, also called Anderson mixing. In the tomographic imaging project, estimation of the geometry parameters can be viewed as a fixed point iteration technique. The research group used three fixed point acceleration schemes in their numerical experiments: 1) Irons-Tuck method, 2) Crossed - second method, 3) Anderson Acceleration. All these acceleration tests were performed and seemed to improve image resolution.&lt;/p>
&lt;p align="center">
&lt;img width="1000" height="500" src="https://user-images.githubusercontent.com/84742324/126915667-fbe60be2-9fa5-4fa1-a7fd-45e910d19ad2.png">
&lt;/p>
&lt;p>From tests performed, the acceleration techniques provided slightly better convergence, with the crossed secant method and Anderson Acceleration performing the best. The Irons-Tuck method converged much better in the angle parameters, but took much longer. The Irons-Tuck image seems to have the least background noise.&lt;/p>
&lt;p align="center">
&lt;img width="1000" height="300" src="https://user-images.githubusercontent.com/84742324/126915702-f55dcd15-4a71-4f4c-bbbd-480734a73394.png">
&lt;/p>
&lt;h2 id="poster-blitz-video">Poster Blitz Video&lt;/h2>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/qdcGe9MKCoI" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen>&lt;/iframe>
&lt;h2 id="more-about-the-team">More About the Team&lt;/h2>
&lt;h3 id="mai-phuong-pham-huynh">Mai Phuong Pham Huynh&lt;/h3>
&lt;p>Mai is a rising junior at Emory University, majoring in Applied Mathematics and Statistics. Her research focus is on numerical analysis and scientific computing. In her free time, she enjoys listening to violin concertos, doing giant Jigsaw puzzles, and searching for vintage vinyl records from the 50s-90s.&lt;/p>
&lt;p align='center'>
&lt;img width=500px, height=auto src="img/IMG_6452.jpeg" class='center'>
&lt;/p>
&lt;h3 id="manuel-santana">Manuel Santana&lt;/h3>
&lt;p>Manuel is a rising senior at Utah State University. His research interests are in all areas of computational and applied math. When not doing math he enjoys European handball, pickleball, raquetball, and cooking with his wife Emily.&lt;/p>
&lt;p align='center'>
&lt;img width=500px, height=auto src="img/linkedInProfile.jpg" class='center'>
&lt;/p>
&lt;h3 id="ana-castillo">Ana Castillo&lt;/h3>
&lt;p>Ana Castillo is in her 10th year teaching. She is certified to teach Spanish and Mathematics grades 6-12 in the states of Texas, Tennessee and Illinois. She is working for Proximity Learning, a company that has allowed her to teach virtually in different school districts in the United States. She has been a participant of the Park City Mathematics Institute Teacher Leadership Program (PCMI TLP) and part of a team that runs a Math Teachers’ Circle in the Rio Grande Valley Area. In her spare time, she enjoys taking long walks, listening to classical music (Vivaldi), spending time with her pets (dogs &amp;amp; donkeys), cooking, reading, writing and traveling.&lt;/p>
&lt;p align="center">
&lt;img width="300" height="400" src="https://user-images.githubusercontent.com/84742324/127031174-4f5717a8-00ee-40df-b342-73f16f2f30c1.jpg">
&lt;/p></description></item><item><title>Efficient Determinant Estimators for Potential Flow Generators</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2021-generative/</link><pubDate>Tue, 14 Dec 2021 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2021-generative/</guid><description>&lt;p>This post was written by Jonathan Valyou, Edward Shao, Lauren Proctor, and Stephen Stern. The project was advised by Dr. Yuanzhe Xi.&lt;/p>
&lt;h2 id="introduction">Introduction&lt;/h2>
&lt;p>Imagine that somehow, you&amp;rsquo;ve recently come to own a massive collection of paintings by famous painters, but none of these works are known to the world. You also happen to be a starving, yet incredibly talented artist, so you wonder about passing off some of your own paintings as part of the collection. You realize that you don&amp;rsquo;t quite have the skills, but you have a friend, a docent at the local museum, that&amp;rsquo;s willing to help you learn.&lt;/p>
&lt;p>After some weeks of thinking the two of you come up with the following plan: each week you present your friend with a randomly chosen painting, either one that you&amp;rsquo;ve painted or one from the real collection. Then they do their best to authenticate it, telling you if they think it is real or a fake, and then you let them know whether or not they were correct.&lt;/p>
&lt;p>Over the course of many weeks you both become better at your jobs because you both get feedback from each other. Each week you try something new, maybe a new technique, brushstroke, or type of paint. So, over time, your skills get better because you learn which combinations work well to trick your friend into thinking your painting is real. On the other hand, each time your friend makes a guess, they learn if they were right or not and learn from their mistakes and successes!&lt;/p>
&lt;p style="font-size: 0.9rem;font-style: italic;">&lt;img style="display: block;" src="https://live.staticflickr.com/7184/6972527877_a36f6004b8_b.jpg" alt="Delaware Art Museum">&lt;a href="https://www.flickr.com/photos/37848681@N05/6972527877">"Delaware Art Museum"&lt;/a>&lt;span> by &lt;a href="https://www.flickr.com/photos/37848681@N05">-Jeffrey-&lt;/a>&lt;/span> is licensed under &lt;a href="https://creativecommons.org/licenses/by-nd/2.0/?ref=ccsearch&amp;atype=html" style="margin-right: 5px;">CC BY-ND 2.0&lt;/a>&lt;a href="https://creativecommons.org/licenses/by-nd/2.0/?ref=ccsearch&amp;atype=html" target="_blank" rel="noopener noreferrer" style="display: inline-block;white-space: none;margin-top: 2px;margin-left: 3px;height: 22px !important;">&lt;img style="height: inherit;margin-right: 3px;display: inline-block;" src="https://search.creativecommons.org/static/img/cc_icon.svg?image_id=0519370b-5525-4215-8237-f933db790ce2" />&lt;img style="height: inherit;margin-right: 3px;display: inline-block;" src="https://search.creativecommons.org/static/img/cc-by_icon.svg" />&lt;img style="height: inherit;margin-right: 3px;display: inline-block;" src="https://search.creativecommons.org/static/img/cc-nd_icon.svg" />&lt;/a>&lt;/p>
&lt;p>Now, at this point you may be wondering why you&amp;rsquo;re reading some hypothetical story about a painter when this is supposed to be about mathematical research, so we&amp;rsquo;ll start getting to the point. The process above is actually the structure of a machine learning technique called Generative Adversarial Network, or GAN for short. In the story, you served the role as a &amp;ldquo;generator&amp;rdquo; by generating new paintings, and your friend serves the role as a &amp;ldquo;discriminator&amp;rdquo; discriminating between the real and fake. In practice GANs work just like the painting scenario, just over the course of what would be equivalent to thousands or millions of weeks. GANs turn out to be incredibly powerful tools which can be used in various ways like deep fakes that you may have seen in the news, or even theorizing new pharmaceutical drugs.&lt;/p>
&lt;h2 id="background">Background&lt;/h2>
&lt;p>Our work does not involve GANs, but it does lie within the the broader field of generative modeling; more specifically our work is in the category of Normalizing Flows. These are also generative models that do not require a discriminator like GANs, but do have some other interesting properties. One essential property is that they are invertible, which has a specific mathematical definition but can more easily be reframed back into our painting example. In this context, invertibility means that there is an exact correspondence between a painting you created and the choices that you made in painting it (e.g. the technique, brushstroke, type of paint, and millions of other small decisions that you made.) This correspondence means that just by looking at your painting I can write down a list of all of the decisions you made. Then if I give you that list I wrote down as a set of instructions you would create the exact same painting!
&lt;figure id="figure-paintings-and-list-of-instructions">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="mainImage" srcset="
/site/cmds-reuret/projects/2021-generative/img/painting_to_list_hu_9593538a721a408f.webp 400w,
/site/cmds-reuret/projects/2021-generative/img/painting_to_list_hu_e2cf503d8cc55931.webp 760w,
/site/cmds-reuret/projects/2021-generative/img/painting_to_list_hu_517d2e4da4d21c19.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2021-generative/img/painting_to_list_hu_9593538a721a408f.webp"
width="611"
height="114"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Paintings and List of Instructions
&lt;/figcaption>&lt;/figure>
It turns out that this setup allows us a way to replace your friend&amp;rsquo;s method of determining if a painting is real or fake. In this context fake means something slightly different: now a fake painting means one that doesn&amp;rsquo;t belong in your collection. If we&amp;rsquo;re given a painting, it&amp;rsquo;s pretty difficult to know whether or not it belongs to your collection by just looking at it, but we can translate that painting into the list of choices that you would have made to paint it. Since you know your style choices pretty well, you can look at that list and tell us how likely it was that you&amp;rsquo;d make that combination of choices. If there&amp;rsquo;s a low probability that you would have made that combination, then the painting probably doesn&amp;rsquo;t belong! This new setup describes how Normalizing Flows work.&lt;/p>
&lt;p>Our work extends &lt;a href="https://arxiv.org/abs/2012.05942" target="_blank" rel="noopener">a paper by Huang et al.&lt;/a> by applying more computationally efficient algorithms for the learning and invertible transformation process they developed. If you are not familiar with neural networks or could use a brush up, we&amp;rsquo;d suggest checking out some or all of these videos by 3 blue 1 brown[link here]. They are a great resource for learning about and getting excited about mathematics and we highly recommend them as a resource. If you are comfortable with neural networks this would be great point to skim the paper that is central to our work, with specific emphasis on section 3.&lt;/p>
&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/videoseries?list=PLZHQObOWTQDNU6R1_67000Dx_ZCJB-3pi" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen>&lt;/iframe>
&lt;h2 id="problem-formulation">Problem Formulation&lt;/h2>
&lt;p>Normalizing flows require an invertible function, like the correspondence between a painting and its list of choices. In general, constructing such a function is not a straightforward process. Huang et al. make use of a property of a special type of functions called convex functions. It turns out that a the gradient of a convex function is invertible, so they propose training an Input Convex Neural Network, which is a convex function, and using its gradient as our normalizing flow.&lt;/p>
&lt;p>In order to transform probabilities from the latent space (our list of choices) and the learning space (paintings) we need to find the determinant of the Jacobian of our normalizing flow. If you are familiar with multivariate calculus, this comes from the change of variable formula. If you are familiar with single variable calculus, it is a higher dimensional analogue of changing dx to du when performing u-substitution. If you were keeping track, this is the Hessian of our convex neural network. In other words, we need the determinant of a matrix that unfortunately takes a lot of computer memory to store, and is impossibly expensive to compute.&lt;/p>
&lt;p>To estimate the determinant without forming the matrix, Huang et al. utilize a tool called a Hutchison Trace Estimator and Stochastic Lanzcos Quadrature (SLQ). This works by deducing (spectral) information about the matrix by randomly sampling a bunch of vectors and measuring how much the matrix changes them through multiplication. This multiplication, called the Hessian vector product, can be performed without explicitly forming the matrix, but is still relatively expensive to compute. This process is required in the forward propagation through the network.&lt;/p>
&lt;p>Another essential part of training neural networks is the loss function. In the first painting story about GANs, the loss function for you was whether or not you fooled your friend, and you learned each week by trying to minimize this function. Your friend learned each week by trying to maximize this function. (This is where the &amp;ldquo;adversarial&amp;rdquo; in the name GAN comes in, you have opposite goals!) The loss function for our work is a bit more complicated, but ultimately is a comparison of true probabilities, and the ones we estimated with the forward transform. This means that classically training our normalizing flow would require backpropogating with the gradient of the forward transform process. Unfortunately this turns out to be highly unstable. Instead, Huang et al. uses a surrogate loss function that combines some algebraic manipulation and another Hutchison Trace Estimator in tandem with an algorithm called the Conjugate Gradient method (CG). This process results in an estimate for the gradient of the loss function required for back propagation. Just like SLQ, CG requires computing a hessian vector product each iteration that it takes to converge on a solution.&lt;/p>
&lt;h2 id="our-contribution">Our Contribution&lt;/h2>
&lt;p>By using information we can deduce from the Stochastic Lanzcos Quadrature, we reduced the number of iterations that CG requires to converge. Specifically, CG approximately solves a linear system of equations, &lt;em>Ax = b&lt;/em>. The number of iterations it takes for the algorithm to converge is related to the condition number of A, the smaller this number is the faster it converges. We apply a pre-conditioner, M, that instead solves the system &lt;em>MAx = Mb&lt;/em>, which is designed to reduce the condition number of MA. Behind the scenes, we can find such a preconditioned because SLQ results in providing the smallest and largest eigenvalues and corresponding eigenvectors of A and the condition number of A is given by the ratio of the largest and smallest eigenvalues.&lt;/p>
&lt;h2 id="more-about-the-team">More About the Team&lt;/h2>
&lt;h2 id="jonathan-valyou">Jonathan Valyou&lt;/h2>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="https://user-images.githubusercontent.com/72425355/127594063-5cf25a7c-3856-4dae-8040-959461793814.jpg" alt="HeadShot" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Jonathan Valyou is a senior at Emory University majoring in Applied Mathematics with a minor in Physics. His primary research areas of interest include scientific computing, data science, and optimization. Aside from this project, he has worked on neural network research. He is currently writing an honors thesis as part of the Emory Mathematics Honors Program. When not tinkering with computational models, he enjoys traveling, listening to classical and pop music, playing the euphonium, and trying new foods.&lt;/p>
&lt;h2 id="lauren-proctor">Lauren Proctor&lt;/h2>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img src="https://avatars.githubusercontent.com/u/64090223?v=4.jpg" alt="HeadShot" loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Lauren Proctor is a junior at the University of Tennessee Majoring in Mathematics and Computer Science. She is a part of the Mathematics Honors program at her university, and has previously worked on natural language processing research. Outside of school, she enjoys hiking, baking, and spending time with friends.&lt;/p></description></item><item><title>Data assimilation for Glacier Modeling</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2022-storm-surge/</link><pubDate>Mon, 27 Jun 2022 11:36:49 -0400</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2022-storm-surge/</guid><description>&lt;p>This post was written by &lt;a href="https://www.linkedin.com/in/emily-corcoran-816278186" target="_blank" rel="noopener">Emily Corcoran&lt;/a>, Hannah Park-Kaufmann, and &lt;a href="https://www.loganknudsen.com/" target="_blank" rel="noopener">Logan Knudsen&lt;/a>. The project was advised by Dr. Talea Mayo. Our team has also created a &lt;a href="img/data_assimilation_for_glacier_modeling.pdf">midterm presentation&lt;/a>, &lt;a href="https://www.youtube.com/watch?v=bGeOZ9G6IOc" target="_blank" rel="noopener">blitz video&lt;/a>, &lt;a href="img/REU_Poster.pdf">poster&lt;/a>, &lt;a href="https://arxiv.org/pdf/2210.02647.pdf" target="_blank" rel="noopener">paper&lt;/a>, and has &lt;a href="https://github.com/hakuupi/StormSurge" target="_blank" rel="noopener">published their code&lt;/a>.&lt;/p>
&lt;h2 id="glaciers">Glaciers&lt;/h2>
&lt;p>Research has shown that climate change will likely impact &lt;a href="https://doi.org/10.3389/fbuil.2020.588049" target="_blank" rel="noopener">storm surge inundation&lt;/a> and make modeling this process more difficult. Sea-level rise caused by climate change plays a part in this impact. To better model sea-level rise, glaciers can be modeled. Marine-Terminating Glaciers have a natural flow towards the ocean, which contributes to sea level rise. By the year 2300, the Antarctic ice sheet is projected to cause up to &lt;a href="https://go.gale.com/ps/i.do?id=GALE%7CA431965879&amp;amp;sid=googleScholar&amp;amp;v=2.1&amp;amp;it=r&amp;amp;linkaccess=abs&amp;amp;issn=00280836&amp;amp;p=HRCA&amp;amp;sw=w&amp;amp;userGroupName=anon%7Eed4bce0c" target="_blank" rel="noopener">3 meters of sea level rise&lt;/a> globally. Due to the severe impacts of glacial melting, modeling changes in ice sheets is an important task. There are challenges to modeling sea level rise, as ice sheet instability leads to significant &lt;a href="https://doi.org/10.1073/pnas.1904822116" target="_blank" rel="noopener">sea-level rise uncertainty&lt;/a>.&lt;/p>
&lt;p align="center">
&lt;img width="500" height="300" src="img/icebergphoto.jpeg">
&lt;/p>
&lt;p align="center">
Image by W. Bulach, used under &lt;a href="https://creativecommons.org/licenses/by-sa/4.0/legalcode">Creative Commons Attribution-Share Alike 4.0 International License&lt;/a>
&lt;/p>
&lt;/p>
&lt;h2 id="modeling-glaciers">Modeling Glaciers&lt;/h2>
&lt;p>Our group is collaborating with &lt;a href="https://iceclimate.eas.gatech.edu/group/" target="_blank" rel="noopener">Dr. Robel&lt;/a>, a glaciologist, climate scientist, and applied mathematician from Georgia Tech, and working with the glacier model described in his &lt;a href="https://doi.org/10.1029/2018JF004709" target="_blank" rel="noopener">2018 paper&lt;/a>. This ice sheet model aims to describe the changes in ice mass of marine-terminating glaciers, which may be impacted over time by climate change.&lt;/p>
&lt;p align="center"> &lt;img width="500" height="300" src="img/glacierdiagram%20(1).png">&lt;/p>
&lt;p align="center">
Image used with permission from &lt;a href="https://doi.org/10.1029/2018JF004709">Dr. Alexander Robel&lt;/a>
&lt;/p>
A glacier can be represented with a simplified box model that has a length $L$, precipitation $P$, and height and flux at the grounding line $h_g$ and $Q_g$. This model is the best approximation for one variable and describes the dominant mode of the glacial system.
&lt;p align="center"> &lt;img width="500" height="300" src="img/boxmodel.png">
&lt;/p>
&lt;!--- Above is the diagram of a box model -->
The two-stage model that our group is using incorporates a nested box into the system. This new box has a thickness, $H$, and an interior flux, $Q$. The change in length and height of the glacier can be described with these differential equations: $$\ \dfrac{dH}{dt}=P-\dfrac{Q_g}{L}-\dfrac{H}{h_gL}(Q-Q_g)$$ $$\ \dfrac{dL}{dt}=\dfrac{1}{h_g}(Q-Q_g)$$
&lt;h2 id="sensitivity-analysis">Sensitivity Analysis&lt;/h2>
&lt;p>Sensitivity analyses study how various sources of uncertainty in a mathematical model contribute to the model&amp;rsquo;s overall uncertainty. This allows us to understand the model better.&lt;/p>
&lt;p>Why do we do this? Why do we care about the uncertainty of a model? And where do the uncertainties even come from?
In this ice sheet model, just like in any model, there are always going to be simplifications, and these lead to uncertainties. We need to have a good idea of which uncertainties matter the most, so that we better know the limits of where our model does a good job of simulating the real world. &lt;br>&lt;/p>
&lt;p>The basic idea is this: We check sensitivity by using different distributions for the input parameters. If the outputs vary significantly, then the output is sensitive to the specification of the input distributions. Hence these should be defined with particular care. We can also look at the sensitivity of the model parameters to inform which parameter we&amp;rsquo;re going to work with in the data assimilation. We want to be working with the most sensitive parameter, because it has the most promise for things we vary later on to matter, in questions like: &amp;ldquo;if your data is from billions of years ago does that matter? Is it important to have your data from the last 60 years?&amp;rdquo; or &amp;ldquo;how much will noise impact the predictions?&amp;rdquo; &lt;br>&lt;/p>
&lt;p>The uncertain model parameters we considered are: initial conditions, sill parameters, and SMB values. For consistency&amp;rsquo;s sake, we vary each parameter by +-10 percent of the nominal values originally given in our model code. Below you can see three graphs, one for each group of parameters varied, for each &amp;ldquo;time vs H(t)&amp;rdquo; (Height of the glacier at time) and &amp;ldquo;time vs L(t)&amp;rdquo; (Length of the glacier at time).&lt;/p>
&lt;p align="center">
&lt;img width="500" height="300" src="img/t_vs_H(t).png">
&lt;/p>
&lt;!--- Above are 3 sensitivity analysis graphs side by side for t vs H(t) -->
&lt;p align="center">
&lt;img width="500" height="300" src="img/t_vs_L(t).png">
&lt;/p>
&lt;!--- Above are 3 sensitivity analysis graphs side by side for t vs L(t) -->
&lt;p>Looking at the distributions, we see that varying initial conditions (Leftmost) seems to produce the greatest spread, but the slopes of the lines there are all very similar. Varying the sill parameters (Middle) produces a lesser spread than varying initial conditions, however there is a greater variation in the slope of the lines. Finally, when varying the smb data (Rightmost) the result actually doesn&amp;rsquo;t change that much and is quite stable. Thus, according to our analysis, the model is the least sensitive to SMB parameters, and between initial and sill parameters judgement varies depending on what you care about more - spread or slope.&lt;/p>
&lt;h2 id="data-assimilation">Data Assimilation&lt;/h2>
&lt;!---Data Assimilation-->
&lt;p>Data assimilation is a method to move models closer to reality using real world observations by readjusting the model state at specified times.&lt;/p>
&lt;p align="center">
&lt;img width="500" height="300" src="img/SEFig.png">
&lt;/p>
&lt;p align="center">
Image used with permission from Dr. Talea Mayo.
&lt;/p>
&lt;p>In this example we have used the ensemble Kalman filter method (ENKF) in order to perform our data assimilation. In basic terms, we initialize an ensemble( or a series of model runs with perturbed initial conditions) and performed data assimilation on each of the ensemble members, then to get our final analysis we took the mean of the ensemble.&lt;/p>
&lt;p align="center">
&lt;img width="500" height="300" src="img/kalmanExample.png">
&lt;/p>
&lt;!--- Above is the Kalman Filter Example -->
&lt;p>The program used to model the glacier behavior and assimilate the data begins with choosing a set of initial conditions. Once the initial conditions are input to the model, which after taking a step using a Runge-Kutta 4th Order Method, is plugged into a Data Assimilation Method. Our main method is ENKF as previously mdentioned. Finally, the analyzed data from the assimilation is output and plugged back into the model. It should be noted that at sometimes the forecast output for the model is the same as the analyzed data.&lt;/p>
&lt;p align="center">
&lt;img width="650" height="300" src="img/ScreenShot2022-07-07at17.21.17.png">
&lt;/p>
&lt;!--- Above is the Kalman Filter Diagram -->
&lt;!---Results-->
&lt;h3 id="square-difference">Square Difference&lt;/h3>
&lt;p>The error measure we use in the best ensemble size and observation scheme is the square difference, $d^2$, which we define $$\ d_t^2 = \left(x_t - x^a_t\right)^2 $$
where $x_t$ is the true state from the truth simulation at time $t$ and $x^a_t$ is the analysis state at time $t$.&lt;/p>
&lt;h3 id="ensemble-size">Ensemble Size&lt;/h3>
&lt;p>In the interest of lowering computational costs, we use the square difference in order to minimize ensemble size while also minimizing error. To do this, we choose an ensemble size, calculate the square difference at each $t$, and then calculated the mean of all these square differences. We ran this calculation for ensembles sizes from 2 up to 75, and found that ensembles of size 7-10 were ideal as they were at the point where the average square difference hovers around the same value.&lt;/p>
&lt;p align="center">
&lt;img width="500" height="400" src="img/Mean_Square_Difference_of_H.png">
&lt;/p>
&lt;p align="center">
&lt;img width="500" height="400" src="img/Mean_Square_Difference_of_L.png">
&lt;/p>
&lt;h3 id="observation-scheme">Observation Scheme&lt;/h3>
&lt;p>We ran the model for various observation schemes to find the best observation scheme, i.e. the times frames and frequencies which can produce a sufficiently small average square difference over the course of the model run. We applied this process to our model and found that for before 1900 the best observation frequency, while still using small number of observations, would be every 19 years for a total of 100 observations. Similarly, for the time frame of 1950-2300 we found that yearly observations for a total of 350 observations is the best frequency.&lt;/p>
&lt;p align="center">
&lt;img width="700" height="400" src="img/oldDates.png">
&lt;/p>
&lt;p align="center">
&lt;img width="700" height="400" src="img/newDates.png">
&lt;/p>
&lt;!---
#### Mean Square Difference 0-1900
|\# Observations | H | L |
|---|---|---|
| 200 | 0.0044054 | 0.0125222 |
| 100 | 0.0087286 | 0.0212535 |
| 50 | 0.0563252 | 0.0613737 |
| 25 | 0.0834296 | 0.0759587 |
| 10 | 0.0952199 | 0.1685447 |
#### Mean Square Difference 1950-2300
|\# Observations | H | L |
|---|---|---|
| 1400 | 0.0025464 | 0.0015147 |
| 700 | 0.0041299 | 0.0019446 |
| 350 | 0.0115712 | 0.0032547 |
| 175 | 0.0175867 | 0.0067435 |
| 88 | 0.0314155 | 0.0172441 |
| 44 | 0.0787454 | 0.0236401 |
| 22 | 0.2535894 | 0.0575641 |
-->
&lt;h3 id="model-runs">Model Runs&lt;/h3>
&lt;p>Using the facts we established in the previous two sections, we ran the model using EnKF for the time frame of 0-2022 in order to project $H$ and $L$ into the future up to the year 2300. The following plots show the results of this experiment, which we will use to help calculate $Q$ and $Q_g$ over time, and in turn use it to calculate sea level rise.&lt;/p>
&lt;p align="center">
&lt;img width="500" height="400" src="img/H(t)projections.png">
&lt;/p>
&lt;p align="center">
&lt;img width="500" height="400" src="img/L(t)projections.png">
&lt;/p>
&lt;h3 id="sea-level-rise">Sea Level Rise&lt;/h3>
&lt;p>Using the formulas for $Q$ and $Q_g$ we can calculate the volume lost across the grounding line
$$\ V_{gz} = W(Q-Q_g)t $$
We then used the to calculate the volume out at all times and add it up to get accumulated volume loss. We then assume that the width of the glacier is 50 km(at least in the case we show here). To convert this to sea level rise, note that 394.67 km$^3$ of ice is equivalent $1$ mm of sea level and get the following projection of sea level rise.&lt;/p>
&lt;p align="center">
&lt;img width="500" height="400" src="img/w50.0km_sea_level_projection.png">
&lt;/p>
&lt;h2 id="next-steps">Next Steps&lt;/h2>
&lt;p>Using data assimilation can help to inform the glacier modelers and glaciologists who collect data about how to collect data in an efficient way. This can help researchers to more efficiently utilize funding and avoid unnecessarily expensive data collection that does not significantly improve glacier models. Data assimilation should be explored within more complicated glacier models, as the model used here is quite simplified. If more research is performed on this technique, it could greatly improve the practice of glacier modeling. Data assimilation can also be used for many geophysical modeling tasks, such as weather forcasting and hurricane storm surge modeling. Going forward, we plan to integrate the output of the glacier model into the ADCIRC hurricane storm surge model to predict the impact of glacier model on storm surge inundation.&lt;/p>
&lt;h2 id="references">References&lt;/h2>
&lt;p>&lt;a href="https://go.gale.com/ps/i.do?id=GALE%7CA431965879&amp;amp;sid=googleScholar&amp;amp;v=2.1&amp;amp;it=r&amp;amp;linkaccess=abs&amp;amp;issn=00280836&amp;amp;p=HRCA&amp;amp;sw=w&amp;amp;userGroupName=anon%7Eed4bce0c" target="_blank" rel="noopener">The long future of Antarctic melting&lt;/a> &lt;br>
&lt;a href="https://doi.org/10.1073/pnas.1904822116" target="_blank" rel="noopener">Marine ice sheet instability amplifies and skews uncertainty in projections of future sea-level rise&lt;/a> &lt;br>
&lt;a href="https://doi.org/10.3389/fbuil.2020.588049" target="_blank" rel="noopener">Projected climate change impacts on hurricane storm surge inundation in the coastal United States&lt;/a> &lt;br>
&lt;a href="https://doi.org/10.1029/2018JF004709" target="_blank" rel="noopener">Response of marine-terminating glaciers to forcing: time scales, sensitivities, instabilities, and stochastic dynamics&lt;/a>&lt;/p>
&lt;h2 id="about-the-team">About the Team&lt;/h2>
&lt;h3 id="emily-corcoran">Emily Corcoran&lt;/h3>
&lt;p>&lt;a href="https://www.linkedin.com/in/emily-corcoran-816278186" target="_blank" rel="noopener">Emily Corcoran&lt;/a> is a junior at New Jersey Institute of Technology, majoring in Mathematical Sciences with a concentration in Applied Statistics and Data Analysis. Before this REU, she has worked as a research assistant in her school&amp;rsquo;s Visual Perception Lab. She is a student in the Albert Dorman Honors College and is an active member of NJIT&amp;rsquo;s school yearbook and Knit &amp;rsquo;n Crochet club. When she is not in class, she can be found reading, listening to music, or attending a local play.&lt;/p>
&lt;h3 id="logan-knudsen">Logan Knudsen&lt;/h3>
&lt;p>&lt;a href="https://www.loganknudsen.com/" target="_blank" rel="noopener">Logan Knudsen&lt;/a> is a senior at Texas A&amp;amp;M University majoring in Mathematics with minors in Oceanography and Meteorology. Before this REU, he has worked doing research on Data Analysis using Benford&amp;rsquo;s Law and as a Teaching Assistant. Logan is currently the President of Texas A&amp;amp;M&amp;rsquo;s Math Club and a member of student radio, KANM. When not in class, he can be found reading, playing the guitar or playing video games with his friends.&lt;/p>
&lt;h3 id="hannah-park-kaufmann">Hannah Park-Kaufmann&lt;/h3>
&lt;p>Hannah Park-Kaufmann is a junior at Bard College and Conservatory, majoring in Mathematics and Piano Performance. Before this REU, she conducted research on Numerical Semigroups and Polyhedra, and on Identifying Universal Traits in Healthy Pianistic Posture using Depth Data. She tutors math in the Bard Prison Initiative (BPI). When not doing math, she can be found playing piano, reading scores, reading literature and/or eating.&lt;/p></description></item><item><title>Fast Training of Implicit Networks with Applications in Inverse Problems</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/</link><pubDate>Mon, 27 Jun 2022 11:36:49 -0400</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/</guid><description>&lt;p>This post was written by Linghai Liu, Shuaicheng Tong, and Lisa Zhao and published with minor edits. The team was advised by Dr. Samy Wu Fung. In addition to this post, the team has also given a &lt;a href="Midterm_Presentation_TeamJFB.pdf">midterm presentation&lt;/a>, filmed a &lt;a href="https://youtu.be/oIwL3E2yULg" target="_blank" rel="noopener">poster blitz video&lt;/a>, created a &lt;a href="REURET_Poster_Team_JFB.pdf">poster&lt;/a>, published &lt;a href="https://github.com/lliu58b/Jacobian-free-Backprop-Implicit-Networks" target="_blank" rel="noopener">code&lt;/a>, and written a &lt;a href="../../publications/liu-et-al-2022/">paper&lt;/a>.&lt;/p>
&lt;h2 id="what-are-inverse-problems">What are Inverse Problems?&lt;/h2>
&lt;p>Inverse problems consist of recovering a signal $x^\ast$ (e.g. an image, a parameter of a PDE, etc.) from indirect, noisy measurements $d$. These problems arise in many applications such as medical imaging, computer vision, geophysical imaging, etc.&lt;/p>
&lt;p>This measurement process is usually modeled as an operator $\mathcal{A}$, satisfying the following equation:
$$ d = \mathcal{A} x^\ast + \boldsymbol{\varepsilon}, $$
where $\mathcal{A}$ is a mapping from signal space $\mathbb{R}^n$ of original images to measurement space $\mathbb{R}^m$. Since our project deals with image deblurring, we have the following variables:&lt;/p>
&lt;ul>
&lt;li>$d \in \mathbb{R}^{n}$: blurred image with noise&lt;/li>
&lt;li>$x^\ast \in \mathbb{R}^{n}$: original image&lt;/li>
&lt;li>$\boldsymbol{\varepsilon} \in \mathbb{R}^{m}$: random &lt;strong>unknown&lt;/strong> noise&lt;/li>
&lt;/ul>
&lt;h2 id="solving-inverse-problems-from-a-classical-approach">Solving Inverse Problems from a Classical Approach&lt;/h2>
&lt;p>Using direct inverse we have:
$$ d = \mathcal{A} x^\ast + \boldsymbol{\varepsilon} \Longrightarrow x^\ast = \mathcal{A}^{-1} d - \mathcal{A}^{-1} \boldsymbol{\varepsilon} $$
However, since $\boldsymbol{\varepsilon}$ is unknown, directly inverting may end up amplifying this noise factor⁠. Because of this noise corruption, the reconstructed image ends up being unrecognizable.&lt;/p>
&lt;p>To better visulaize this, we have the following set of pictures:&lt;/p>
&lt;p>
&lt;figure id="figure-original-image">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Original Image" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/inverse1_hu_9a31eb4fd5e9a7ee.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse1_hu_b5b0cf880240cb01.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse1_hu_b55a9ada36ce7f14.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/inverse1_hu_9a31eb4fd5e9a7ee.webp"
width="231"
height="231"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Original Image
&lt;/figcaption>&lt;/figure>
&lt;figure id="figure-blurred-noisy-image">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Blurred Noisy Image" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/inverse2_hu_ebddb48c5044fed5.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse2_hu_3aec6aad3315a09f.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse2_hu_ace3300dd198ca2b.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/inverse2_hu_ebddb48c5044fed5.webp"
width="231"
height="231"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Blurred Noisy Image
&lt;/figcaption>&lt;/figure>
&lt;figure id="figure-direct-inverse">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Direct Inverse" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/inverse3_hu_7481675a2a9a075a.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse3_hu_ea7a24caefdcfb4e.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse3_hu_5c52c617251f33ce.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/inverse3_hu_7481675a2a9a075a.webp"
width="231"
height="231"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Direct Inverse
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;p>In order to minimize the noise factor, we want to formulate a regularized optimization problem.&lt;/p>
&lt;p>We essentially want to find the minimium distance between the reconstructed image and the observed blurred image, plus a regularizer $R(x)$.&lt;/p>
&lt;p>This regularizer is chosen based on prior knowledge of the data; this can often lead to inaccuracies—meaning the reconstructed image will be a bit blurry.&lt;/p>
&lt;p>For example, using a gradient descent scheme where we handpick a regularizer to help stabilize the reconstruction, we have the following set of pictures:&lt;/p>
&lt;p>
&lt;figure id="figure-original-image">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Original Image" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/inverse1_hu_9a31eb4fd5e9a7ee.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse1_hu_b5b0cf880240cb01.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse1_hu_b55a9ada36ce7f14.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/inverse1_hu_9a31eb4fd5e9a7ee.webp"
width="231"
height="231"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Original Image
&lt;/figcaption>&lt;/figure>
&lt;figure id="figure-blurred-noisy-image">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Blurred Noisy Image" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/inverse2_hu_ebddb48c5044fed5.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse2_hu_3aec6aad3315a09f.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/inverse2_hu_ace3300dd198ca2b.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/inverse2_hu_ebddb48c5044fed5.webp"
width="231"
height="231"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Blurred Noisy Image
&lt;/figcaption>&lt;/figure>
&lt;figure id="figure-gradient-descent">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Gradient Descent" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/gd4_hu_9b4381ba6a2098a.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/gd4_hu_728281e70438475b.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/gd4_hu_378e72fcb6750716.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/gd4_hu_9b4381ba6a2098a.webp"
width="231"
height="231"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Gradient Descent
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;p>We see that the reconstructed image using gradient descent is a huge improvement from direct inverse; however, there are still blurry areas we can improve on. In search of a better method, we turn towards implicit learning.&lt;/p>
&lt;h2 id="implicit-deep-learning">Implicit Deep Learning&lt;/h2>
&lt;p>The issue with the classical approach is that the regularizer is chosen heristically. To combat this, our approach now is to utilize data to &lt;strong>learn&lt;/strong> and &lt;strong>train&lt;/strong> the regularizer.&lt;/p>
&lt;p>To do this we mimic gradient descent, but replace the gradient of the regularizer, $\lambda \nabla_x R$, with a trainable network.&lt;/p>
&lt;p>However, this creates some problems concerning memory cost and the number of layers, $K$, in our neural network. The memory grows linearly as $K$—chosen heuristically—increases.&lt;/p>
&lt;p>With implicit deep learning we send $K \to \infty$ until we find a fixed point of a single layer $T_\Theta(\cdot)$.&lt;/p>
&lt;h2 id="implicit-backpropagation">Implicit Backpropagation&lt;/h2>
&lt;p>Suppose now we have found a fixed point $x^\ast$ for a single layer. Then, $$ x^\ast = T_\Theta (x^\ast) $$&lt;/p>
&lt;p>Using implicit differentiation on the equation above we have,&lt;/p>
&lt;p>$$\frac{d x^\ast}{d \Theta} = \left( I - \frac{d T_\Theta (x^\ast)}{d x^\ast}\right)^{-1} \frac{\partial T_\Theta (x^\ast)}{\partial \Theta}$$&lt;/p>
&lt;p>However, solving this is very expensive because of the inverse term.&lt;/p>
&lt;p>To circumvent this issue, we use a recently proposed method called &lt;a href="https://arxiv.org/abs/2103.12803" target="_blank" rel="noopener">Jacobian-Free Backpropagation&lt;/a>.&lt;/p>
&lt;h2 id="jacobian-free-backpropagation-jfb">Jacobian-Free Backpropagation (JFB)&lt;/h2>
&lt;p>The goal of JFB is to alleviate memory requirement and avoid high computational cost in implicit networks.&lt;/p>
&lt;p>The key idea is to replace the problematic Jacobian $$\left( I - \frac{d T_\Theta (x^\ast)}{d x^\ast}\right)$$ with the identity matrix $I$.&lt;/p>
&lt;p>For a comparison, if we were to calculate the true gradient using implicit networks we have the following equation:&lt;/p>
&lt;p>$$\nabla_\Theta \ell = \frac{d \ell}{d x^\ast} \left( I - \frac{d T_\Theta (x^\ast)}{d x^\ast}\right) ^{-1} \frac{\partial T_\Theta (x^\ast)}{\partial \Theta}$$&lt;/p>
&lt;p>Using JFB to approximate the gradient we only need to solve:
$$p_\Theta = \frac{d \ell}{d x^\ast} \frac{\partial T_\Theta (x^\ast)}{\partial \Theta}$$
which is a descent direction for the loss $\ell$.&lt;/p>
&lt;p>Utilizing JFB, we avoid computing the Jacobian term. As a result, implicit networks are trained faster and more easily implemented.&lt;/p>
&lt;p>Note: the JFB approach relies on a set of conditions to be true:&lt;/p>
&lt;ul>
&lt;li>$T_\Theta$ is contraction mapping with Lipschitz constant $\gamma$&lt;/li>
&lt;li>$T_\Theta$ is continuously differentiable w.r.t. $\Theta$&lt;/li>
&lt;li>$M := \frac{\partial T_\Theta}{\partial \Theta}$ has full column rank&lt;/li>
&lt;li>$M$ is well-conditioned, i.e., $\kappa (M^T M) &amp;lt; \frac{1}{\gamma}$&lt;/li>
&lt;/ul>
&lt;h2 id="numerical-experiments">Numerical Experiments&lt;/h2>
&lt;p>In our project we used the CelebA dataset, which consist of annotated celebrity faces. The images are categorized into various sections based on specific features that the celebrities have. For example, whether or not they have bangs, wear glasses, have a pointy nose, etc.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="CelebA" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/celebA_hu_47e6f403fa15e501.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/celebA_hu_cb15b8c9354af515.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/celebA_hu_fe3eec365473bd10.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/celebA_hu_47e6f403fa15e501.webp"
width="760"
height="517"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h2 id="results">Results&lt;/h2>
&lt;p>The results are as follows
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="loss plot" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/loss_plot-1_hu_cbed88ce59921863.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/loss_plot-1_hu_12c1d557739f59d6.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/loss_plot-1_hu_ac9e68c8c5c40697.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/loss_plot-1_hu_cbed88ce59921863.webp"
width="760"
height="545"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
From the graph we see that the loss is decreasing as the number of epochs increases. An epoch is one complete pass of the entire dataset through our algorithm.
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="results" srcset="
/site/cmds-reuret/projects/2022-implicit/imgs/truth_36-1_hu_9d4791c3ac39e077.webp 400w,
/site/cmds-reuret/projects/2022-implicit/imgs/truth_36-1_hu_9c158a254c0fa600.webp 760w,
/site/cmds-reuret/projects/2022-implicit/imgs/truth_36-1_hu_d13e724c304ac897.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-implicit/imgs/truth_36-1_hu_9d4791c3ac39e077.webp"
width="760"
height="324"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
Note: Two metrics are commonly used for assessing the quality of reconstructed images: the peak-signal-to-noise ratio (PSNR, a positive number, best at $+\infty$) and the structural similarity index measure (SSIM, also positive, best at $1$).&lt;/p>
&lt;h2 id="acknowledgements">Acknowledgements&lt;/h2>
&lt;p>We sincerely thank the guidance of our mentor, Dr. Samy Wu Fung, and other mentors at Emory University for the opportunity.&lt;/p>
&lt;h2 id="more-about-the-team">More About the Team&lt;/h2>
&lt;p>&lt;strong>Linghai Liu&lt;/strong> is a rising senior at Brown University, double concentrating in applied mathematics - computer science and mathematics. His main interests lie at the intersection of statistical theory, machine learning, and optimization. Outside of work, he enjoys reading novels and watching animes.&lt;/p>
&lt;p>&lt;strong>Shuaicheng Tong&lt;/strong> is a rising junior at the University of California, Los Angeles, majoring in applied mathematics and minoring in statistics. He is interested in optimization and machine learning. He volunteers at the UCLA Statistics Club where he tutors mathematics and statistics. Outside of school, he enjoys working out, hiking, and watching Star Wars shows.&lt;/p>
&lt;p>&lt;strong>Lisa Zhao&lt;/strong> is a rising sophomore at the University of California, Berkeley, double majoring in statistics and economics. She is interested in learning about how statistics is used as a powerful tool in finance. Outside of work, she enjoys swimming, drawing, and watching TV shows.&lt;/p></description></item><item><title>Learning Ordinary Differential Equations from Data</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2022-learn-ode/</link><pubDate>Mon, 27 Jun 2022 11:36:49 -0400</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2022-learn-ode/</guid><description>&lt;!-- --- -->
&lt;!-- # Emory REU Learn ODE: -->
&lt;!--Starting new section-->
&lt;!-- --- -->
&lt;p>This post was written by &lt;a href="https://ehayes75.github.io/" target="_blank" rel="noopener">Emma Hayes&lt;/a>, &lt;a href="https://mathheider.github.io/" target="_blank" rel="noopener">Mathias Heider&lt;/a>, and &lt;a href="https://cvanty.github.io/" target="_blank" rel="noopener">Carrie Vanty&lt;/a> and published with minor edits. The team was advised by Dr. Deepanshu Verma.&lt;/p>
&lt;p>In addition to this post, the team has also created slides for a &lt;a href="pdfs/presentation.pdf">midterm presentation&lt;/a>, a &lt;a href="">poster blitz video&lt;/a>, and a &lt;a href="pdfs/poster.pdf">poster&lt;/a>.&lt;/p>
&lt;h2 id="project-overview">Project Overview:&lt;/h2>
&lt;p>Imagine a spring mass system (Figure 1). What if you wanted to find the location of the mass at any given time point? In order to find this information, you must first understand the dynamics of the system. A spring mass system is an example of simple harmonic motion where total energy is conserved. This means that you can model the dynamics using a Hamiltonian Ordinary Differential Equation, which has the quality of energy conservation. To solve our problem, we use neural networks utilizing Hamiltonians in the forward propagation to predict our coordinates.&lt;/p>
&lt;p align="center">
&lt;img src=images/Simple_harmonic_oscillator.gif>
&lt;p align = "center">
Figure 1 - Spring Mass by Oleg Alexandrov (public domain)
&lt;/p>
&lt;p>Our project aims to compute the value of the Hamiltonian for any given time and set of initial conditions using Hamiltonian Inspired neural networks. We will first introduce the mathematical background of our project and the novel technique we implemented for our forward propagation. Results will then be presented and analyzed. Lastly, we will discuss how our project expands upon both Ruthotto &lt;a href="https://arxiv.org/abs/1705.03341" target="_blank" rel="noopener">3&lt;/a> and Greydanus &lt;a href="https://arxiv.org/abs/1906.01563" target="_blank" rel="noopener">6&lt;/a> papers.&lt;/p>
&lt;h2 id="background">Background:&lt;/h2>
&lt;p>Often, neural networks are thought of as a black box, where the actual inner-workings are not the main focus. However, since we are mathematicians, we want to understand how the network functions in order to best optimize it. For this reason, we began by looking at why and how ODEs were first used in neural networks. Ordinary differential equations were first used in Residual Neural Networks due to the similarity between the forward propagation equation and discretization of an ordinary differential equation. The only difference is multiplication of the step size, which we denote as $\mathbf{h}$.
$$Y_{j+1} = Y_j + \mathbf{h}\sigma(Y_j K_j + b_j)$$
In the context of our residual neural network, the ODE as forward propagation means that for each layer of the network, we will move one time step forward in the discretization of our network ODE. The weights and biases, $K$ and $b$, may change in between the layers depending on the given values in the network ODE. The output of our network is the Hamiltonian value at the given time, and from that we are able to approximate position and velocity values for the mass.
When estimating coordinates of a Hamiltonian system, or the value of the Hamiltonian itself, the Hamiltonian relationships are important for forward propagation. Hamiltonians intrinsically conserve energy, meaning the network is better able to learn conservation laws and predict examples with energy conservation. Without considering these relationships, it would be much more difficult for the network to learn conservation, which can cause a buildup of error. In many studies (&lt;a href="https://arxiv.org/abs/1705.03341" target="_blank" rel="noopener">3&lt;/a>, &lt;a href="https://arxiv.org/abs/1906.01563" target="_blank" rel="noopener">6&lt;/a>), they have found that by using the Hamiltonian equations in Hamiltonian data sets, error has decreased. We plan to investigate this further and find which discretization methods and algorithms will perform the best.&lt;/p>
&lt;h2 id="learning-hamiltonians-from-data">Learning Hamiltonians from Data&lt;/h2>
&lt;p>To learn about Hamiltonian dynamics from data, we use neural networks. Specifically a modified version of the Residual Neural Network (RNN), which we call a Hamiltonian Inspired Neural Network (HINN) drawn from &lt;a href="https://arxiv.org/abs/1705.03341" target="_blank" rel="noopener">3&lt;/a>. To create this HINN, we primarily used 2 packages - PyTorch and hessQuik. The difference between our HINN, and the traditional RNN and Ruthotto \textit{et al}’ HINN, is in our forward propagation method and how we input values into our MSE loss function. The forward propagation uses the autograd feature to calculate both $\frac{\partial H_{\theta}}{\partial \mathbf{p}}$ and $\frac{\partial H_{\theta}}{\partial \mathbf{q}}$, where $\theta$ are the network parameters we wish to optimize and $H_{\theta}$ is our network output. We then use these values in discretizing $\mathbf{p_{\theta}}$ and $\mathbf{q_{\theta}}$.
$$\mathbf{p_{\theta+1}} = \mathbf{p_{\theta}} + h\frac{\partial H_{\theta}}{\partial \mathbf{q}} $$
$$\mathbf{q_{\theta+1}} = \mathbf{q_{\theta}} - h\frac{\partial H_{\theta}}{\partial \mathbf{p}}$$&lt;/p>
&lt;p>The new values are then plugged into our MSE loss function. Using these techniques we &lt;a href="https://github.com/mathheider/Learn-ODEs-HINNs" target="_blank" rel="noopener">created two HINNs&lt;/a> for two different examples, those being the Simple Spring Mass System and the Two Body Problem.&lt;/p>
&lt;p>Results:&lt;/p>
&lt;p align="center">
&lt;img src=images/findingNemo.png width = "800">
&lt;p align = "center">
Figure 2 - Spring Mass System 1. Training Loss Graph 2. Learned Position Values over True Position Values 3. Relationship of Learned p and q over ground truth
&lt;/p>
&lt;p align="center">
&lt;img src=images/2Body.png width = "800">
&lt;p align = "center">
Figure 3 - Two Body Problem 1. Training Loss Graph 2. Learned Trajectories Graph over True Trajectories 3. Learned Energy over True Energy 4. Learned Position over True Position.
&lt;/p>
&lt;h2 id="whats-next">What&amp;rsquo;s Next&lt;/h2>
&lt;p>Currently our results look promising as our learned values closely match the ground truth values. Moving foward, we would like to add more complexity to our current examples: instead of fixed time steps, we would like to attempt variable time steps. We would also like to take on the Three Body Problem, which unlike our current examples has no analytical solution.&lt;/p>
&lt;h2 id="information-about-us">Information about Us&lt;/h2>
&lt;p>Mathias Heider is a rising senior at the University of Delaware, majoring in Computer Science and Mathematics and Economics. His interest are in machine learning and data science specifically when it relates to dataset with bioinformatics applications. Outside of class, Mathias likes to ski, hangout with friends, and play video games&lt;/p>
&lt;p>Carrie Vanty is a rising senior at Middlebury College, majoring in mathematics. She loves working with ordinary differential equations in applied math. Outside of school, Carrie likes to ski, hike, and craft.&lt;/p>
&lt;p>Emma Hayes is a rising junior at Carnegie Mellon University, majoring in Computational and Applied Mathematics. She enjoys learning about new topics in mathematics and incorporating computer science into her work. Outside of class, Emma likes hiking, baking, and making art.&lt;/p>
&lt;h2 id="references">References&lt;/h2>
&lt;p>[1] &lt;a href="https://en.wikipedia.org/wiki/Effective_mass_%28spring%E2%80%93mass_system%29#/media/File:Simple_harmonic_oscillator.gif" target="_blank" rel="noopener">https://en.wikipedia.org/wiki/Effective_mass_(spring%E2%80%93mass_system)#/media/File:Simple_harmonic_oscillator.gif&lt;/a>&lt;/p>
&lt;p>[2] &lt;a href="https://arxiv.org/abs/1512.03385" target="_blank" rel="noopener">Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016.&lt;/a>&lt;/p>
&lt;p>[3] &lt;a href="https://arxiv.org/abs/1705.03341" target="_blank" rel="noopener">Eldad Haber and Lars Ruthotto. Stable architectures for deep neural networks. Inverse Problems, 34(1):014004, 22, 2018.&lt;/a>&lt;/p>
&lt;p>[4] &lt;a href="https://doi.org/10.21105/joss.02104" target="_blank" rel="noopener">Brian de Silva, Kathleen Champion, Markus Quade, Jean-Christophe Loiseau, J. Kutz, and Steven Brunton. Pysindy: A python package for the sparse identification of nonlinear dynamical systems from data. Journal of Open Source Software, 5(49):2104, 2020.&lt;/a>&lt;/p>
&lt;p>[5] &lt;a href="https://doi.org/10.21105/joss.03994" target="_blank" rel="noopener">Alan A. Kaptanoglu, Brian M. de Silva, Urban Fasel, Kadierdan Kaheman, Andy J. Goldschmidt, Jared Callaham, Charles B. Delahunt, Zachary G. Nicolaou, Kathleen Champion, Jean-Christophe Loiseau, J. Nathan Kutz, and Steven L. Brunton. Pysindy: A comprehensive python package for robust sparse system identification. Journal of Open Source Software, 7(69):3994, 2022.&lt;/a>&lt;/p>
&lt;p>[6] &lt;a href="https://arxiv.org/abs/1906.01563" target="_blank" rel="noopener">Sam Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian neural networks. ArXiv, abs/1906.01563, 2019.&lt;/a>&lt;/p></description></item><item><title>Low-Precision Algorithms for Image Processing</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2022-mixed-precision/</link><pubDate>Mon, 27 Jun 2022 11:36:49 -0400</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2022-mixed-precision/</guid><description>&lt;p>This post was written by &lt;a href="https://github.com/kristinagxy" target="_blank" rel="noopener">Xiaoyun Gong&lt;/a>, &lt;a href="https://github.com/RileyCYZ" target="_blank" rel="noopener">Yizhou Chen&lt;/a>, and &lt;a href="https://github.com/zoejix" target="_blank" rel="noopener">Xiang Ji&lt;/a> and published with minor edits. The team was advised by Dr. &lt;a href="https://github.com/jnagy1" target="_blank" rel="noopener">James Nagy&lt;/a>. In addition to this post, the team has also created slides for a &lt;a href="REUmidterm_presentation.pdf">midterm presentation&lt;/a>, a &lt;a href="https://www.dropbox.com/s/139l0u7zi6eloao/mixed-precision.mp4?dl=0" target="_blank" rel="noopener">poster blitz&lt;/a> video, &lt;a href="https://github.com/kristinagxy/REU_code" target="_blank" rel="noopener">code&lt;/a>, and a &lt;a href="">paper&lt;/a>.&lt;/p>
&lt;h2 id="in-one-sentence">In One Sentence:&lt;/h2>
&lt;p>Our group works on experimenting with iterative methods for solving inverse problems at different precision levels.&lt;/p>
&lt;h2 id="background-why-low-precision">Background: Why Low Precision?&lt;/h2>
&lt;p>What is the most important aspect for an excellent gaming experience? A lot of people would answer real-time! Everyone wants their games to be fast, and it is always a bummer that the screen freezes during a critical combat. This is why we are investigating low precision arithmetic: to decrease the computation time and speed things up.&lt;/p>
&lt;p>Nowadays, most computer systems operate on double precision (64-bit) arithmetic. However, if we decrease the number of bits for each number to 16 bits or even lower, the processing time can be much significantly reduced, although the benefit comes at the cost of a loss of accuracy.&lt;/p>
&lt;p align="center">
&lt;img src="img/IEEE.png" alt="drawing" width="500"/>
&lt;/p>
&lt;p>
&lt;em>Using tensor cores for mixed-precision scientific computing.
Oct-2021. url: https://developer.nvidia.com/blog/tensor-cores-mixed-precision-scientific-computing/
&lt;/em>
&lt;/p >
&lt;h2 id="simulating-low-precision">Simulating Low Precision&lt;/h2>
&lt;h3 id="matlab-function-chop">Matlab function chop&lt;/h3>
&lt;p>To simulate low precision arithmetic on our 64-bit computers, we have imported a MATLAB package called &lt;a href="https://www.mathworks.com/matlabcentral/fileexchange/70651-chop?s_tid=mwa_osa_a" target="_blank" rel="noopener">chop&lt;/a>. The toolbox allows us to explore single precision, half precision, and other customized formats. Each input needs to be transformed, but the real work comes from chopping each operation. The code below is a toy example of how to calculate $x + y \times z$ in half precision with chop.&lt;/p>
&lt;p align="center">
&lt;img src="img/chop_overview.png" alt="drawing" width="300"/>
&lt;/p>
&lt;h3 id="blocking">Blocking&lt;/h3>
&lt;p>When the number is being chopped from double precision to half precision, a lot of bits are dropped (from 64 bits to 16 bits). This would certainly cause a level of inaccuracy, so in order to reduce the errors, a method called blocking is used. Blocking is the same as breaking a large operation into smaller chunks, where each is computed independently and the result is then summed.&lt;/p>
&lt;p>We compute the inner product for each precision and block size for 20 times and calculate the average. The errors are calculated as the differences between the result of using the chopped version of inner product function and the built-in function in matlab.&lt;/p>
&lt;p align="center">
&lt;img src="img/blocksize.png" alt="draw" width="700"/>
&lt;/p>
On the left-hand side of the graph where the size of the vector is 1000, the errors of half precision are the largest because it has the least bits. If we take a closer look at only half precision, we get the graph on the right with different vector sizes. The errors decrease sharply when the blocking method is introduced. However, larger block sizes do not necessarily mean lower errors, as the graph suggested: the errors increase again as the block size keeps growing. That is because when the block size is large, it's the same as doing no blocking at all. For example, for a size-500 vector, once the block size reaches 500, it just means putting the whole vector into the first block, the same as when blocking is not introduced. Therefore, the line becomes flat from 500. We use 256 as our default block size in our codes because the matrix dimension is rather large in our problem.
&lt;h2 id="inverse-problems-and-iterative-methods">Inverse Problems and Iterative Methods&lt;/h2>
&lt;p>Inverse problems are problems where our goal is to find the internal or hidden information (inputs) from outside measurements (results). The internal data can be approximated by iterative methods, a repeating implementation of the same system of equations, with variables getting updated each round, hoping the value generated can be closer to the true value we desire each term.&lt;/p>
&lt;h3 id="conjugate-gradient-method">Conjugate Gradient Method&lt;/h3>
&lt;p>The &lt;a href="https://www.cs.cmu.edu/~quake-papers/painless-conjugate-gradient.pdf" target="_blank" rel="noopener">Conjugate Gradient algorithm&lt;/a> (CG) aims to solve the linear system Ax = b where A is SPD (symmetric and positive definite),transforming the problem of finding solution to an optimization problem where we want to minimize $\phi(x)=\frac{1}{2}x^{T}Ax-x^{T}b$. This can be easily seen from $\nabla \phi (x) = 0$ -&amp;gt; $Ax-b=0$.&lt;/p>
&lt;p>In each step, the method provides us with a search direction and a step-length so that the error of this iteration is A-orthogonal to the search direction of the previous iteration. Eventually, it will converge to the minimal point. The CGLS algorithm is the least-squares version of the CG method, applied to the normal equation A&lt;sup>T&lt;/sup>Ax = A&lt;sup>T&lt;/sup>b. However, CGLS requires computing inner products, which can overflow for large-scale problems in low precision.&lt;/p>
&lt;h3 id="chebyshev-semi-iterative-method">Chebyshev Semi-Iterative Method&lt;/h3>
&lt;p>The Chebyshev Semi-Iterative (CS) Method requires no inner product computation, which is great because inner products can cause overflow easily in low precision. But there is always the trade-off! The CS method requires the user to have an idea of the range of the matrix A&amp;rsquo;s eigenvalues. The result given by CS is a linear combination of all solutions in each iteration, and the weights are obtained from the Chebyshev polynomial, which has the favorable property to ensure that the result obtained in each iteration of CS is smaller than an upper bound.&lt;/p>
&lt;h2 id="experiment">Experiment&lt;/h2>
&lt;h3 id="ir-tools">IR Tools&lt;/h3>
&lt;p>We modify the CGLS method in the &lt;a href="https://github.com/jnagy1/IRtools.git" target="_blank" rel="noopener">IRtool&lt;/a> package in Matlab so that it can operate in lower precision, and we use two test problems in the same package to investigate how the method performs at lower precision, mainly half precision.&lt;/p>
&lt;h3 id="image-deblurring-using-cgls">Image Deblurring Using CGLS&lt;/h3>
&lt;p>First, we use our modified version of CGLS without regularization to solve the image deblurring problem. In this application, we solve Ax = b, where b is an observed blurred image, A is a matrix that models the blurring operation, and x is the desired clean image. We didn’t add any noise to b in the problem of Ax = b at the beginning, and the graphs are demonstrated below.&lt;/p>
&lt;p align="center">
&lt;img src="img/blur no noise.png" alt="draw" width="600"/>
&lt;/p>
&lt;p>We use our modified version of CGLS for single and half precision, and the graph in single precision is similar to the graph in double precision. However, for half precision, the background is not the same as that in the double-precision or single-precision graph; it contains more artifacts.&lt;/p>
&lt;p>We also plot the error norms of the solution at each iteration using different precision levels.&lt;/p>
&lt;p align="center">
&lt;img src="img/enrm blur no noise.png" alt="draw" width="600"/>
&lt;/p>
&lt;p>From the graph, all three error norms overlap from the beginning until around the 20th iteration, where the half-precision errors begin to deviate from those in single and double precision. The difference is due to the round-off errors of half precision, which add up and take over. Besides, the error norms for half precision terminates at the 28th iteration because overflow of inner products causes NaNs (Not a Number) to be computed during the iteration.&lt;/p>
&lt;p>After investigating the idealized situations where there is no noise in the observed image, we then apply our code to problems that contains additive random noise to see how it is likely to perform in real life. That is, we try to compute x from the observed image b = Ax + noise.&lt;/p>
&lt;p>For half precision, with 0.1% noise, the picture looks almost the same as the one that contains no noise. However, if the noise level is increased to 1%, the background has substantially more artifacts, while the middle object is still identifiable. Noise has taken over the black background but not the satellite yet. Eventually, the whole image is flushed with the noisy artifacts with 10% noise; the picture no longer contains any meaningful information. Notice that the results below are generated using x from the best iteration, that is the iteration with the smallest error norms, not from the last iteration.&lt;/p>
&lt;p align="center">
&lt;img src="img/Best cgls fp16 64 m noise.png" alt="draw" width="700"/>
&lt;/p>
&lt;p>Now we turn our attention to the error norm, the difference between the original image and the one our algorithm generates at each iteration. When 0.1% noise is added, as the number of iterations goes up, the error norm reduces significantly across all three formats. Intriguingly, for images with 1% or 10% noise, the best reconstruction is not the last iteration but somewhere along the middle (it’s around the 50th iteration for 1% and 10th for 10%). The reason behind the phenomenon is that while we are transforming the output image, b, the blended noise also gets inverted along each iteration. Eventually, the random data accumulate and dominate the solution at some point. We are showing the results where the error norm is the smallest to see what is the best possible solution we can compute. However, in reality the true x is not known, meaning we don&amp;rsquo;t know the error norms, so we can only show results from the last iteration, not from the best iteration.&lt;/p>
&lt;p align="center">
&lt;img src="img/3.png" alt="draw" width="700"/>
&lt;/p>
&lt;!--
### Tomography Reconstruction Using CGLS
Below is the result of CGLS on the tomography reconstruction test problem at different precision levels. For double and single precision, the reconstruction is doing well, yet for fp16, we start to get this completely blue picture from the first iteration caused by overflow of Inf/-Inf.
&lt;p align="center">
&lt;img src="img/tomo_plot.png" alt="draw" width="800"/>
&lt;/p>
Our solution to this issue is to rescale A and b by dividing both by 100. And we get the result below:
&lt;p align="center">
&lt;img src="img/tomo_rescale.png" alt="draw" width="300"/>
&lt;/p>
which is still blurry but at least the shape is visible :)
Then we add noise to the right-hand side b, and plot the error norms below:
&lt;p align="center">
&lt;img src="img/error_tomo.png" alt="draw" width="800"/>
&lt;/p>
As in the image deblurring problem, the error norms first decrease and then increase. In the cases with noise, this is mainly because noise starts to take over in the later part of the iteration. However, we still see the same behavior in the noise-free test problems at half precision, which is because the truncation errors accumulate as the iteration goes on.
-->
&lt;h3 id="image-deblurring-using-cs">Image Deblurring Using CS&lt;/h3>
&lt;p>In order to prevent the occurrence of overflow, we experiment with the CS algorithm (where no inner products are needed) and use chop for lower precision. Tikhonov regularization is applied to CS after we find out that the algorithm performs poorly due to the close-to-zero singular values of A when it&amp;rsquo;s ill-conditioned. Now we are solving:
$$\min_{x} {||Ax-b||_2^2+\lambda^2||x||_2^2}$$ where $\lambda$ is a parameter that needs to be chosen. Here we show experiments for the case with 10% noise, and we use $\lambda$ = 0.199.&lt;/p>
&lt;p>From the graph below, it is clear that even with 10% noise, the half-precision image looks very similar to that in double precision, better than what we have using CGLS (results from the last iteration).&lt;/p>
&lt;p align="center">
&lt;img src="img/cs_reg_0.1_blur.png" alt="draw" width="500"/>
&lt;/p>
For the image deblurring problem, we further comfirm the similarity by plotting the error norms.
&lt;p align="center">
&lt;img src="img/cs_reg_0.1_blur_Enrm.png" alt="draw" width="300"/>
&lt;/p>
We can see that the error norms of the three precision levels overlap, illustrating that the result in half precision is close to that in double precision.
&lt;!--
### Tomography Reconstruction Using CS
The result for the tomography reconstruction problem using CS is showed below.
&lt;p align="center">
&lt;img src="img/cs_reg_0.1_tomo.png" alt="draw" width="500"/>
&lt;/p>
Although there is no inner product in CS, we still have overflow in half precision, which is because the matrix A is too large. Therefore, we rescale A and b by diving both of by 100 again, and the algorithm successfully runs to the end without generating NaNs. The image in half precision is as good as its counterpart in double precision, displaying clear boundaries and backgrounds. The error norms also overlap among the three precision levels.
-->
&lt;h3 id="image-deblurring-using-cgls-with-regularization">Image Deblurring Using CGLS with regularization&lt;/h3>
&lt;p>To fairly compare CGLS and CS, we add Tikhonov regularization to CGLS and run the test problem again. The diagrams are listed below.&lt;/p>
&lt;p align="center">
&lt;img src="img/cg_reg_blur_64_0.1_m.png" alt="draw" width="500"/>
&lt;/p>
&lt;p align="center">
&lt;img src="img/cg_reg_m_0.1_blur_Enrm.png" alt="draw" width="300"/>
&lt;/p>
The diagram for half precision looks much better than that produced by CGLS without regularization. However, difference still presents between half and double precision in the background. At the end of the iteration for half precision, the error norms still increase again. If we zoom in the graph of the error norms for CS and CGLS with regularization, we can see that the error norms at half precision converge for CS but increases rapidly for CGLS, suggesting that for half precision, CS is a better choice, especially when the noise level is high. When the noise level is close to zero, CS becomes susceptible because of the accumulation of round-off errors. However, for double precision, the CGLS method with regularization is clearly more stable. Therefore, CGLS with regularization is more suitable for double precision.
&lt;p align="center">
&lt;img src="img/Zoom in of Enrm for cs and cg.png" alt="draw" width="500"/>
&lt;/p>
We performed similar experiments with an image reconstruction problem from tomography; see our paper for further details!
&lt;h2 id="more-about-us">More about us&lt;/h2>
&lt;h3 id="yizhou-chen">Yizhou Chen&lt;/h3>
&lt;p align="center">
&lt;img src="img/bioRiley.png" alt="draw" width="300"/>
&lt;/p>
Hi! My name is Yizhou, but I go by Riley as well. I'm a rising junior at Emory University, double majoring in Applied Math and Physics. My interest is in computational Math &amp; computational Physics. Outside work I enjoy watching sitcoms and my favourite one is Frasier! I am an animal person and I have a toy poodle who's nine years old. I also love cycling and hiking.
&lt;h3 id="xiaoyun-gong">Xiaoyun Gong&lt;/h3>
&lt;p align="center">
&lt;img src="img/gxy.jpeg" alt="drawing" width="300"/>
&lt;/p>
Hello I am Xiaoyun Gong. I am a rising senior majoring in Applied Mathematics and Statistics. I am interested in math and I also enjoy coding!! In my free time I like watching anime (most recent favorite is Made in Abyss) and drawing. I like sweet food and I am a cat person. 🐱
&lt;h3 id="xiang-ji">Xiang Ji&lt;/h3>
&lt;p align="center">
&lt;img src="img/Zoe_bio.png" alt="drawing" width="300"/>
&lt;/p>
Hi, I am Xiang Ji, but you can also call me Zoe. I am a rising junior at Emory University who is double majoring in applied mathematics and statistics and art history. My research interest is in computational mathematics and image processing. I enjoy going to art museums and watching movies. It's quite fun doing research this summer!</description></item><item><title>Model-based approaches to neuronal network firing and its subsequent validation with a previously recorded in-vivo dataset</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/</link><pubDate>Mon, 27 Jun 2022 11:36:49 -0400</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/</guid><description>&lt;p>The research featured in this blog post was performed by Carly Ferrell, Qile Jiang, and Olivia Leu, and the team was advised by Dr. Michael Caiola.
This blog post was written by Carly Ferrell and published with minor edits. In addition to this post, the team has also given a &lt;a href="https://mstate-my.sharepoint.com/:b:/g/personal/cgf115_msstate_edu/EXa7BOlzUEhOui75S-m7CDABtmi4HFTbWnezklPHVSGadA?e=4XxeOW">midterm presentation&lt;/a>, made a &lt;a href="https://www.youtube.com/watch?v=2uBVgNFRpqI">poster blitz video&lt;/a>, and created a &lt;a href="https://mstate-my.sharepoint.com/:b:/g/personal/cgf115_msstate_edu/ERy6qqUmC7NPg20btpNQ0acBQOcWI51R2BewUUZX9tmvDQ?e=CpfiyS">poster&lt;/a>. They are currently working on a paper with the aim to publish it in an academic journal.&lt;/p>
&lt;h1 id="mathematical-modeling-of-healthy-and-parkinsonian-firing-patterns-in-the-primate-thalamocortical-motor-circuit">Mathematical Modeling of Healthy and Parkinsonian Firing Patterns in the Primate Thalamocortical Motor Circuit&lt;/h1>
&lt;p>Parkinson&amp;rsquo;s disease (PD) is a slowly progressing neuro-degenerative disease featuring impaired motor symptoms such as bradykinesia, muscular rigidity, and resting tremors.&lt;sup>1&lt;/sup> In industrialized countries, PD affects 0.3% of all people and 1% of people over age 60.&lt;sup>6&lt;/sup> The basal ganglia, motor thalamus, and motor cortex are three main components of the brain&amp;rsquo;s motor circuit and are responsible for movement planning and execution; movement disorders such as PD can develop when the typical activity of this circuit is disrupted.&lt;sup>2,3&lt;/sup> Specifically, PD is associated with the loss of dopaminergic neurons and altered neuronal oscillations in the beta-band (13-30 Hz).&lt;sup>4&lt;/sup> Other projects, such as the &lt;a href="https://www.worldscientific.com/doi/epdf/10.1142/S0129065718500211">2019 paper by M. Caiola and M. Holmes&lt;/a>, have investigated the changes in the basal ganglia neuronal activity from a mathematical modeling perspective, but little research has been done on the parkinsonism-associated changes in the areas of the thalamus and cortex which are involved in the motor circuit.&lt;sup>5&lt;/sup> We employ a mathematical model to investigate network connection changes within the thalamocortical motor ciruit to better understand the transition from healthy to parkinsonian states in the brain.&lt;/p>
&lt;h1 id="firing-rate-model">Firing Rate Model&lt;/h1>
&lt;p>We choose to use a firing rate model to describe our system. This approach can successfully represent networks, since each unit in the model can represent a population of neurons receiving input (average firing rates) from other neuron populations.&lt;/p>
&lt;p>A simplified circuit diagram of the thalamocortical motor circuit network is shown below, and provides the neuroscience basis for our model. The rounded squares each represent a population of neurons, which are connected by either excitatory (arrow-tipped lines) or inhibitory (circle-tipped lines) synaptic weights. The green circle represents the interneuron population of the thalamus.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Thalamocortical Loop Model" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/Thalamocortical_hu_852d92d8dd0195e9.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/Thalamocortical_hu_7d1f0863fa6c1ff7.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/Thalamocortical_hu_9fa63b6c7eaf99ad.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/Thalamocortical_hu_852d92d8dd0195e9.webp"
width="476"
height="490"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;table>
&lt;tr>
&lt;td>GPi (&lt;em>y&lt;/em>&lt;sub>1&lt;/sub>)&lt;/td>
&lt;td>globus pallidus internal&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>TC (&lt;em>y&lt;/em>&lt;sub>2&lt;/sub>)&lt;/td>
&lt;td>thalamocortical neurons&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>CT5 (&lt;em>y&lt;/em>&lt;sub>3&lt;/sub>)&lt;/td>
&lt;td>corticothalamic layer 5&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>CT6 (&lt;em>y&lt;/em>&lt;sub>4&lt;/sub>)&lt;/td>
&lt;td>corticothalamic layer 6&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>RTN (&lt;em>y&lt;/em>&lt;sub>5&lt;/sub>)&lt;/td>
&lt;td>thalamic reticular nucleus&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>IN (&lt;em>&amp;gamma;&lt;/em>)&lt;/td>
&lt;td>thalamic interneuron population&lt;/td>
&lt;/tr>
&lt;/table>
&lt;p>Treating the interneuron population as a &amp;ldquo;relay,&amp;rdquo; &lt;em>γ&lt;/em>, we can establish the following system of equations:&lt;/p>
&lt;p>&lt;em>τ&lt;sub>1&lt;/sub>y&amp;rsquo;&lt;sub>1&lt;/sub>&lt;/em> = −&lt;em>y&lt;sub>1&lt;/sub>&lt;/em> + &lt;em>f&lt;sub>1&lt;/sub>&lt;/em>(&lt;em>β&lt;sub>1&lt;/sub>&lt;/em> + &lt;em>h&lt;/em>)&lt;/p>
&lt;p>&lt;em>τ&lt;sub>2&lt;/sub>y&amp;rsquo;&lt;sub>2&lt;/sub>&lt;/em> = −&lt;em>y&lt;sub>2&lt;/sub>&lt;/em> + &lt;em>f&lt;sub>2&lt;/sub>&lt;/em>(&lt;em>w&lt;sub>12&lt;/sub>y&lt;sub>1&lt;/sub>&lt;/em> + &lt;em>w&lt;sub>32&lt;/sub>y&lt;sub>3&lt;/sub>&lt;/em> + &lt;em>w&lt;sub>42&lt;/sub>y&lt;sub>4&lt;/sub>&lt;/em> − &lt;em>w&lt;sub>52&lt;/sub>y&lt;sub>5&lt;/sub>&lt;/em> + &lt;em>γ&lt;/em> + &lt;em>b&lt;sub>2&lt;/sub>&lt;/em>)&lt;/p>
&lt;p>&lt;em>τ&lt;sub>3&lt;/sub>y&amp;rsquo;&lt;sub>3&lt;/sub>&lt;/em> = −&lt;em>y&lt;sub>3&lt;/sub>&lt;/em> + &lt;em>f&lt;sub>3&lt;/sub>&lt;/em>(&lt;em>w&lt;sub>23&lt;/sub>y&lt;sub>2&lt;/sub>&lt;/em> + &lt;em>w&lt;sub>43&lt;/sub>y&lt;sub>4&lt;/sub>&lt;/em> + &lt;em>b&lt;sub>3&lt;/sub>&lt;/em>)&lt;/p>
&lt;p>&lt;em>τ&lt;sub>4&lt;/sub>y&amp;rsquo;&lt;sub>4&lt;/sub>&lt;/em> = −&lt;em>y&lt;sub>4&lt;/sub>&lt;/em> + &lt;em>f&lt;sub>4&lt;/sub>&lt;/em>(&lt;em>w&lt;sub>34&lt;/sub>y&lt;sub>3&lt;/sub>&lt;/em> + &lt;em>b&lt;sub>4&lt;/sub>&lt;/em>)&lt;/p>
&lt;p>&lt;em>τ&lt;sub>5&lt;/sub>y&amp;rsquo;&lt;sub>5&lt;/sub>&lt;/em> = −&lt;em>y&lt;sub>5&lt;/sub>&lt;/em> + &lt;em>f&lt;sub>5&lt;/sub>&lt;/em>(&lt;em>w&lt;sub>45&lt;/sub>y&lt;sub>4&lt;/sub>&lt;/em> + &lt;em>b&lt;sub>5&lt;/sub>&lt;/em>)&lt;/p>
&lt;p>γ = −&lt;em>w&lt;sub>62&lt;/sub>&lt;/em>(−&lt;em>w&lt;sub>16&lt;/sub>y&lt;sub>1&lt;/sub>&lt;/em> − &lt;em>w&lt;sub>56&lt;/sub>y&lt;sub>5&lt;/sub>&lt;/em> + &lt;em>w&lt;sub>46&lt;/sub>y&lt;sub>4&lt;/sub>&lt;/em> + &lt;em>b&lt;sub>6&lt;/sub>&lt;/em>)&lt;/p>
&lt;table>
&lt;tr>
&lt;td>&lt;em>y&lt;sub>i&lt;/sub>&lt;/em>&lt;/td>
&lt;td>average neuronal population firing rate&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;em>w&lt;sub>jk&lt;/sub>&lt;/em>&lt;/td>
&lt;td>weight of the connection between populations &lt;em>j&lt;/em> and &lt;em>k&lt;/em>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;em>h&lt;/em>&lt;/td>
&lt;td>constant basal ganglia input&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;em>&amp;tau;&lt;sub>i&lt;/sub>&lt;/em>&lt;/td>
&lt;td>membrane time constant&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;em>f&lt;sub>i&lt;/sub>&lt;/em>&lt;/td>
&lt;td>activation function&lt;/td>
&lt;/tr>
&lt;/table>
&lt;p>Note that &lt;em>w&lt;sub>23&lt;/sub>&lt;/em> represents the difference between the excitatory and inhibitory inputs from TC to CT5. Note also that &lt;em>w&lt;sub>jk&lt;/sub>&lt;/em> &amp;gt; 0 and τ&lt;sub>&lt;em>i&lt;/em>&lt;/sub> &amp;gt; 0.&lt;/p>
&lt;p>This can be represented with vectors and matrices as:&lt;/p>
&lt;p>&lt;em>T&lt;strong>y&amp;rsquo;&lt;/strong>&lt;/em> = −&lt;em>&lt;strong>y&lt;/strong>&lt;/em> + &lt;em>&lt;strong>F&lt;/strong>&lt;/em>(&lt;em>&lt;strong>x&lt;/strong>&lt;/em>) ⟹ &lt;em>T&lt;strong>y&amp;rsquo;&lt;/strong>&lt;/em> = &lt;em>A&lt;strong>y&lt;/strong>&lt;/em> + &lt;em>&lt;strong>B&lt;/strong>&lt;/em>&lt;/p>
&lt;h1 id="activation-function-selection">Activation Function Selection&lt;/h1>
&lt;p>Neurons traditionally respond to inputs sigmoidally.&lt;sup>7,8,9&lt;/sup> However, this model creates a nonlinear system of equations for which it is impossible to solve for eigenvalues analytically. In order to attain eigenvalues and be able to comment on the behavior of the model as a whole, we must establish a simpler activation function that still manages to approximate experimental neuron discharge behavior.&lt;sup>5&lt;/sup> A piecewise linear (PWL) activation function is ideal in our case, as it allows us to break down a complex system into linear pieces which can be solved and manipulated:&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Piecewise Linear Activation Function" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/pwl_act_func_hu_a2c7911e6f3567c3.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/pwl_act_func_hu_1977598ecdd8a66d.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/pwl_act_func_hu_896e0bf40d0274fc.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/pwl_act_func_hu_a2c7911e6f3567c3.webp"
width="560"
height="321"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>We can break down this system into 3&lt;sup>5&lt;/sup> = 243 distinct linear regions in space, each with its own steady state (fixed point in space which the solution tends to as time increases). Out of these 243 regions, only the region in which each activation function is between 0 spikes/sec and its maximum firing rate contains a physiologically realistic steady state, further denoted as the middle region (outlined in green in the diagram below).&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Accuracy of Piecewise Linear Activation Function to the Sigmoidal Activation Function" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/pws_sig_hu_e35ed553029b038d.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/pws_sig_hu_67cf72c45d9e8f70.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/pws_sig_hu_8780fcbee87505c8.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/pws_sig_hu_e35ed553029b038d.webp"
width="716"
height="585"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h1 id="data-matching">Data Matching&lt;/h1>
&lt;p>This semi-linear firing rate model has a number of constant values that we must locate in experimental data and incorporate, namely the baseline firing rates, maximum firing rates, and membrane time constants for each neuron population involved in our simplified motor circuit model. We were able to find values for these parameters through literature review, although some required that we make estimates informed by information from areas of the brain that behave similarly or data on these parameters from mice, rats, or cats. However, there does not seem to be data that documents the baseline firing rate for the thalamic interneuron population in the primate brain. Given our uncertainty about the true baseline firing rate value for the primate thalamic interneuron population, we decided to create two models, one with the low and one with the high baseline. The parameter values are shown in the table below:&lt;/p>
&lt;table>
&lt;tr>
&lt;td>Neuron Population&lt;/td>
&lt;td>&lt;em>b&lt;sub>i&lt;/sub>&lt;/em>&lt;/td>
&lt;td>&lt;em>M&lt;sub>i&lt;/sub>&lt;/em>&lt;/td>
&lt;td>&lt;em>&amp;tau;&lt;sub>i&lt;/sub>&lt;/em>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>GPi (&lt;em>y&lt;/em>&lt;sub>1&lt;/sub>)&lt;/td>
&lt;td>55 Hz&lt;sup>5,10,11,12,13&lt;/sup>&lt;/td>
&lt;td>200 Hz&lt;sup>14&lt;/sup>&lt;/td>
&lt;td>8 ms&lt;sup>11,15,16&lt;/sup>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>TC (&lt;em>y&lt;/em>&lt;sub>2&lt;/sub>)&lt;/td>
&lt;td>18.5 Hz&lt;sup>17&lt;/sup>&lt;/td>
&lt;td>300 Hz&lt;/td>
&lt;td>25 ms&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>CT5 (&lt;em>y&lt;/em>&lt;sub>3&lt;/sub>)&lt;/td>
&lt;td>7.25 Hz&lt;sup>18&lt;/sup>&lt;/td>
&lt;td>200 Hz&lt;/td>
&lt;td>20 ms&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>CT6 (&lt;em>y&lt;/em>&lt;sub>4&lt;/sub>)&lt;/td>
&lt;td>7.25 Hz&lt;sup>18&lt;/sup>&lt;/td>
&lt;td>200 Hz&lt;/td>
&lt;td>15 ms&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>RTN (&lt;em>y&lt;/em>&lt;sub>5&lt;/sub>)&lt;/td>
&lt;td>25 Hz&lt;sup>17&lt;/sup>&lt;/td>
&lt;td>500 Hz&lt;sup>17&lt;/sup>&lt;/td>
&lt;td>16.51 ms&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>IN (&lt;em>&amp;gamma;&lt;/em>)&lt;/td>
&lt;td>Low: 6 Hz&lt;sup>19&lt;/sup> &lt;br> High: 22.7 Hz&lt;sup>20&lt;/sup>&lt;/td>
&lt;td>N/A&lt;/td>
&lt;td>N/A&lt;/td>
&lt;/tr>
&lt;/table>
&lt;h1 id="stability-and-steady-state-conditions">Stability and Steady State Conditions&lt;/h1>
&lt;p>No matter the disease state of our model, the neurons should not be at a state of maximal firing or absent firing for an extended period of time. Additionally, in Parkinsonian solutions, we should expect oscillations of firing rates. Thus the following must hold:&lt;/p>
&lt;ol>
&lt;li>Middle region contains its own steady state, and trajectories must not stabilize in another region.&lt;/li>
&lt;li>&lt;strong>Healthy:&lt;/strong> Middle region is stable ⟶ trajectories are thus forced to stabilize in the middle region, making the system globally asymptotically stable. &lt;br> &lt;strong>Parkinsonian:&lt;/strong> Middle region is unstable ⟶ trajectories are thus forced to oscillate around the middle region, forming a globally stable limit cycle.&lt;/li>
&lt;/ol>
&lt;p>To determine stability, the PWL activation function allows us to solve for the eigenvalues of each of the 243 regions explicitly. We found 3 possible cases:&lt;/p>
&lt;ol>
&lt;li>The region is stable regardless of weights.&lt;/li>
&lt;li>The region&amp;rsquo;s stability is conditional on weight values.&lt;/li>
&lt;li>The region (including the middle region) has eigenvalues that cannot be solved for analytically. Therefore, we used the &lt;strong>Routh-Hurwitz Stability Criterion&lt;/strong> (RH) to derive 3 stability conditions.&lt;/li>
&lt;/ol>
&lt;h1 id="weight-search">Weight Search&lt;/h1>
&lt;p>The current literature does not specify the baseline firing rate for the interneuron population, &lt;em>b&lt;sub>6&lt;/sub>&lt;/em>, so we took two estimates: &lt;em>b&lt;sub>6&lt;/sub>&lt;/em> = 6 for the low estimate, and &lt;em>b&lt;sub>6&lt;/sub>&lt;/em> = 22.7 for the high estimate.&lt;/p>
&lt;p>Comparing our data to the predicted values our model outputted, we were able to minimize the sum of squared error between the two and find a healthy solution for both the low and the high estimates of &lt;em>b&lt;sub>6&lt;/sub>&lt;/em>. The outputs for the low &lt;em>b&lt;/em>&lt;sub>6&lt;/sub> are shown below:&lt;/p>
&lt;p>Healthy Solution:&lt;/p>
&lt;table>
&lt;tr>
&lt;td>&lt;em>w&lt;sub>12&lt;/sub>&lt;/em> = 1.520384442&lt;/td>
&lt;td>&lt;em>w&lt;sub>16&lt;/sub>&lt;/em> = 1.621278311&lt;/td>
&lt;td>&lt;em>w&lt;sub>23&lt;/sub>&lt;/em> = 0.4962387866&lt;/td>
&lt;td>&lt;em>w&lt;sub>32&lt;/sub>&lt;/em> = 1.117631687&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;em>w&lt;sub>34&lt;/sub>&lt;/em> = 0.1540248925&lt;/td>
&lt;td>&lt;em>w&lt;sub>42&lt;/sub>&lt;/em> = 1.217895798&lt;/td>
&lt;td>&lt;em>w&lt;sub>43&lt;/sub>&lt;/em> = 0.0672671083&lt;/td>
&lt;td>&lt;em>w&lt;sub>45&lt;/sub>&lt;/em> = 1.542582263&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;em>w&lt;sub>46&lt;/sub>&lt;/em> = 9.049109867&lt;/td>
&lt;td>&lt;em>w&lt;sub>52&lt;/sub>&lt;/em> = 4.5350845&lt;/td>
&lt;td>&lt;em>w&lt;sub>56&lt;/sub>&lt;/em> = 0.3330689302&lt;/td>
&lt;td>&lt;em>w&lt;sub>62&lt;/sub>&lt;/em> = 7.127373038&lt;/td>
&lt;/tr>
&lt;/table>
&lt;p>Below is shown the firing rate outputs using these weights.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Healthy Low FR" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/healthylow_FR_hu_4bc13fd989cccc01.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/healthylow_FR_hu_9a4ef69cf81db261.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/healthylow_FR_hu_5deb8fb18adfc778.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/healthylow_FR_hu_4bc13fd989cccc01.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Here, all firing rates tend toward a specifc value as time increases, so they are stable solutions.&lt;/p>
&lt;p>Parkinsonian Solution:&lt;/p>
&lt;table>
&lt;tr>
&lt;td>&lt;em>w&lt;sub>12&lt;/sub>&lt;/em> = 1.520384442&lt;/td>
&lt;td>&lt;em>w&lt;sub>16&lt;/sub>&lt;/em> = 1.621278311&lt;/td>
&lt;td>&lt;em>w&lt;sub>23&lt;/sub>&lt;/em> = 0.8691494663&lt;/td>
&lt;td>&lt;em>w&lt;sub>32&lt;/sub>&lt;/em> = 0.5043792396&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;em>w&lt;sub>34&lt;/sub>&lt;/em> = 0.1540248925&lt;/td>
&lt;td>&lt;em>w&lt;sub>42&lt;/sub>&lt;/em> = 1.217895798&lt;/td>
&lt;td>&lt;em>w&lt;sub>43&lt;/sub>&lt;/em> = 0.0672671083&lt;/td>
&lt;td>&lt;em>w&lt;sub>45&lt;/sub>&lt;/em> = 1.542582263&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;em>w&lt;sub>46&lt;/sub>&lt;/em> = 9.049109867&lt;/td>
&lt;td>&lt;em>w&lt;sub>52&lt;/sub>&lt;/em> = 4.5350845&lt;/td>
&lt;td>&lt;em>w&lt;sub>56&lt;/sub>&lt;/em> = 0.3330689302&lt;/td>
&lt;td>&lt;em>w&lt;sub>62&lt;/sub>&lt;/em> = 7.127373038&lt;/td>
&lt;/tr>
&lt;/table>
&lt;p>Below is shown the firing rate outputs using these weights.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="PD Low FR" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/PDlow_FR_hu_3901f5742cad9208.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/PDlow_FR_hu_324d3694a6af1a2b.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/PDlow_FR_hu_b6e2149b70181679.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/PDlow_FR_hu_3901f5742cad9208.webp"
width="760"
height="570"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Here, several firing rates oscillate, indicating a limit cycle solution.&lt;/p>
&lt;h1 id="weight-space">Weight Space&lt;/h1>
&lt;p>We were interested in the role of the thalamus in parkinsonian dysfunction, so we explored the relationship between &lt;em>w&lt;/em>&lt;sub>23&lt;/sub> and &lt;em>w&lt;/em>&lt;sub>32&lt;/sub>, which represent the excitatory and inhibitory connections between TC and CT5. Forcing all correlating weights to be equal in healthy and parkinsonian solutions except &lt;em>w&lt;/em>&lt;sub>23&lt;/sub> and &lt;em>w&lt;/em>&lt;sub>32&lt;/sub>, we found a healthy solution that could be forced into a parkinsonian state by only altering &lt;em>w&lt;/em>&lt;sub>23&lt;/sub> and &lt;em>w&lt;/em>&lt;sub>32&lt;/sub>. In a parkinsonian solution, at least one of the Routh-Hurwitz stability conditions must be broken. We examined which condition or combination of conditions is broken when the system moves from a healthy to a parkinsonian state for different values of &lt;em>w&lt;/em>&lt;sub>23&lt;/sub> and &lt;em>w&lt;/em>&lt;sub>32&lt;/sub>. The region plot for the low &lt;em>b&lt;/em>&lt;sub>6&lt;/sub> is shown below.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Region Plot Low b6" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/regplot_web_hu_2797a3b0d82b8c27.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/regplot_web_hu_fd8e16e665bbcac5.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/regplot_web_hu_f14f74bcda396b03.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/regplot_web_hu_2797a3b0d82b8c27.webp"
width="760"
height="522"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h1 id="conclusions-and-future-directions">Conclusions and Future Directions&lt;/h1>
&lt;ul>
&lt;li>Our model can represent the average firing rates of healthy and parkinsonian states in the thalamocortical motor circuit.&lt;/li>
&lt;li>We established stability and steady state conditions for the system to be healthy or parkinsonian.&lt;/li>
&lt;li>We have found multiple sets of weights that both satisfy the conditions and match the neuronal firing patterns in our &lt;em>in-vivo&lt;/em> primate dataset.&lt;/li>
&lt;li>We discovered that changing only the connection strength between TC and CT5 can force the system from a healthy to a parkinsonian state.&lt;/li>
&lt;li>&lt;strong>Next steps:&lt;/strong>&lt;/li> &lt;ul>
&lt;li>Examine the transition from both healthy to parkinsonian and parkinsonian to healthy.&lt;/li>
&lt;li>Explore methods of biologically validating our model by pharmacologically manipulating the weights of motor circuit network connections.&lt;/li> &lt;/ul>
&lt;/ul>
&lt;h1 id="more-about-the-team">More About the Team&lt;/h1>
&lt;ol>
&lt;li>&lt;strong>Carly Ferrell&lt;/strong> is a rising senior at Mississippi State University majoring in mathematics and minoring in statistics and music with a concentration in voice. She is interested in utilizing her skills in applied mathematics and statistcs to research music, specifically music theory and sight singing. Outside class, she enjoys reading, dancing, singing, and composing music.&lt;/li>
&lt;/ol>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Carly Picture" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/carly_pic_hu_6fa205f5d60d588f.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/carly_pic_hu_cc34f20642edab50.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/carly_pic_hu_d1bb65a9596bb3c0.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/carly_pic_hu_6fa205f5d60d588f.webp"
width="760"
height="507"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;ol start="2">
&lt;li>&lt;strong> Qile Jiang&lt;/strong> is a rising junior at Brown University majoring in Applied Mathematics. His primary research area is in applied dynamical systems, but he also has a keen interest in pure math topics such as algebra. Outside of school, he spends his time training boxing, painting, and going to operas and classical concerts.&lt;/li>
&lt;/ol>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Qile Picture" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/qile_pic_hu_159de25a42f1d826.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/qile_pic_hu_b6942e985028bfab.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/qile_pic_hu_b836125051211188.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/qile_pic_hu_159de25a42f1d826.webp"
width="722"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;ol start="3">
&lt;li>&lt;strong>Margaret Olivia Leu&lt;/strong> is a junior at Pomona College double majoring in mathematics and politics. She is interested in working on ways to use mathematics as a tool in the fields of politics and social justice work, and hopes to pursue a career that combines these two interests. Outside academics, she enjoys crocheting, cooking, and listening to music.&lt;/li>
&lt;/ol>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Olivia Picture" srcset="
/site/cmds-reuret/projects/2022-neural-net-firing/olivia_pic_hu_9a673c612e6a4861.webp 400w,
/site/cmds-reuret/projects/2022-neural-net-firing/olivia_pic_hu_9c7f236ce7ba2a57.webp 760w,
/site/cmds-reuret/projects/2022-neural-net-firing/olivia_pic_hu_e6dadbd87978971d.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2022-neural-net-firing/olivia_pic_hu_9a673c612e6a4861.webp"
width="754"
height="760"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;h1 id="references">References&lt;/h1>
&lt;ol>
&lt;li>Sveinbjornsdottir, S. (2016).The clinical symptoms of Parkinson&amp;rsquo;s disease. &lt;em>Journal of Neurochemistry, 139&lt;/em>(1), 318-324. &lt;a href="https://doi.org/10.1111/jnc.13691" target="_blank" rel="noopener">https://doi.org/10.1111/jnc.13691&lt;/a>.&lt;/li>
&lt;li>DeLong, M. R., &amp;amp; Wichmann, T. (2007). Circuits and circuit disorders of the basal ganglia. &lt;em>Archives of Neurology, 64&lt;/em>(1), 20–24. &lt;a href="https://doi.org/10.1001/archneur.64.1.20" target="_blank" rel="noopener">https://doi.org/10.1001/archneur.64.1.20&lt;/a>.&lt;/li>
&lt;li>Alexander, G. E., DeLong, M.R., &amp;amp; Strick, P.L. (1986). Parallel Organization of functionally segregated circuits linking basal ganglia and cortex. &lt;em>Annual Review of Neuroscience, 9&lt;/em>(1), 357-381. &lt;a href="https://doi.org/10.1146/annurev.ne.09.030186.002041" target="_blank" rel="noopener">https://doi.org/10.1146/annurev.ne.09.030186.002041&lt;/a>&lt;/li>
&lt;li>Galvan, A., Devergnas, A., &amp;amp; Wichmann, T. (2015). Alterations in neuronal activity in basal ganglia-thalamocortical circuits in the parkinsonian state. &lt;em>Frontiers in Neuroanatomy, 9&lt;/em>, 5. &lt;a href="https://doi.org/10.3389/fnana.2015.00005" target="_blank" rel="noopener">https://doi.org/10.3389/fnana.2015.00005&lt;/a>.&lt;/li>
&lt;li>Caiola, M., &amp;amp; Holmes, M. H. (2019). Model and analysis for the onset of parkinsonian firing patterns in a simplified basal ganglia. &lt;em>International Journal of Neural Systems, 29&lt;/em>(1). &lt;a href="https://doi.org/10.1142/S0129065718500211" target="_blank" rel="noopener">https://doi.org/10.1142/S0129065718500211&lt;/a>.&lt;/li>
&lt;li>de Lau, L. M. L., &amp;amp; Breteler, M. M. B. (2006). Epidemiology of Parkinson&amp;rsquo;s disease. &lt;em>The Lancet Neurology, 5&lt;/em>(6), 525-535. &lt;a href="https://doi.org/10.1016/S1474-4422%2806%2970471-9" target="_blank" rel="noopener">https://doi.org/10.1016/S1474-4422(06)70471-9&lt;/a>.&lt;/li>
&lt;li>Rall, W. (1955). Experimental monosynaptic input-output relations in the mammalian spinal cord. &lt;em>Journal of Cellular and Comparative Physiology, 46&lt;/em>(3), 413–437. &lt;a href="https://doi.org/10.1002/jcp.1030460303" target="_blank" rel="noopener">https://doi.org/10.1002/jcp.1030460303&lt;/a>&lt;/li>
&lt;li>Wilson, C. J., &amp;amp; Bevan, M. D. (2011). Intrinsic dynamics and synaptic inputs control the activity patterns of subthalamic nucleus neurons in health and in Parkinson’s disease. &lt;em>Neuroscience, 198&lt;/em>, 54–68. &lt;a href="https://doi.org/10.1016/j.neuroscience.2011.06.049" target="_blank" rel="noopener">https://doi.org/10.1016/j.neuroscience.2011.06.049&lt;/a>&lt;/li>
&lt;li>Nambu, A., &amp;amp; Llinaś, R. (1994). Electrophysiology of globus pallidus neurons in vitro. &lt;em>Journal of Neurophysiology, 72&lt;/em>(3), 1127–1139. &lt;a href="https://doi.org/10.1152/jn.1994.72.3.1127" target="_blank" rel="noopener">https://doi.org/10.1152/jn.1994.72.3.1127&lt;/a>&lt;/li>
&lt;li>Kita, H., Tachibana, Y., Nambu, A., &amp;amp; Chiken, S. (2005). Balance of Monosynaptic Excitatory and Disynaptic Inhibitory Responses of the Globus Pallidus Induced after Stimulation of the Subthalamic Nucleus in the Monkey. &lt;em>Journal of Neuroscience, 25&lt;/em>(38), 8611–8619. &lt;a href="https://doi.org/10.1523/JNEUROSCI.1719-05.2005" target="_blank" rel="noopener">https://doi.org/10.1523/JNEUROSCI.1719-05.2005&lt;/a>&lt;/li>
&lt;li>Hikosaka, O. (2007). GABAergic output of the basal ganglia. &lt;em>Progress in Brain Research, 160&lt;/em>, 209–226. &lt;a href="https://doi.org/10.1016/S0079-6123%2806%2960012-5" target="_blank" rel="noopener">https://doi.org/10.1016/S0079-6123(06)60012-5&lt;/a>&lt;/li>
&lt;li>Wichmann, T., Bergman, H., Starr, P. A., Subramanian, T., Watts, R. L., &amp;amp; DeLong, M. R. (1999). Comparison of MPTP-induced changes in spontaneous neuronal discharge in the internal pallidal segment and in the substantia nigra pars reticulata in primates. &lt;em>Experimental Brain Research, 125&lt;/em>(4), 397–409. &lt;a href="https://doi.org/10.1007/s002210050696" target="_blank" rel="noopener">https://doi.org/10.1007/s002210050696&lt;/a>&lt;/li>
&lt;li>Bergman, H., Wichmann, T., Karmon, B., &amp;amp; DeLong, M. R. (1994). The primate subthalamic nucleus. II. Neuronal activity in the MPTP model of parkinsonism. &lt;em>Journal of Neurophysiology, 72&lt;/em>(2), 507–520. &lt;a href="https://doi.org/10.1152/jn.1994.72.2.507" target="_blank" rel="noopener">https://doi.org/10.1152/jn.1994.72.2.507&lt;/a>&lt;/li>
&lt;li>Hashimoto, T., Elder, C. M., Okun, M. S., Patrick, S. K., &amp;amp; Vitek, J. L. (2003). Stimulation of the Subthalamic Nucleus Changes the Firing Pattern of Pallidal Neurons. &lt;em>The Journal of Neuroscience, 23&lt;/em>(5), 1916–1923. &lt;a href="https://doi.org/10.1523/JNEUROSCI.23-05-01916.2003" target="_blank" rel="noopener">https://doi.org/10.1523/JNEUROSCI.23-05-01916.2003&lt;/a>&lt;/li>
&lt;li>Nakanishi, H., Tamura, A., Kawai, K., &amp;amp; Yamamoto, K. (1997). Electrophysiological studies of rat substantia nigra neurons in an in vitro slice preparation after middle cerebral artery occlusion. &lt;em>Neuroscience, 77&lt;/em>(4), 1021–1028. &lt;a href="https://doi.org/10.1016/s0306-4522%2896%2900555-6" target="_blank" rel="noopener">https://doi.org/10.1016/s0306-4522(96)00555-6&lt;/a>&lt;/li>
&lt;li>Nambu, A. (2007). Globus pallidus internal segment. &lt;em>Progress in Brain Research, 160&lt;/em>, 135–150. &lt;a href="https://doi.org/10.1016/S0079-6123%2806%2960008-3" target="_blank" rel="noopener">https://doi.org/10.1016/S0079-6123(06)60008-3&lt;/a>&lt;/li>
&lt;li>van Albada, S. J., &amp;amp; Robinson, P. A. (2009). Mean-field modeling of the basal ganglia-thalamocortical system. I Firing rates in healthy and parkinsonian states. &lt;em>Journal of Theoretical Biology, 257&lt;/em>(4), 642–663. &lt;a href="https://doi.org/10.1016/j.jtbi.2008.12.018" target="_blank" rel="noopener">https://doi.org/10.1016/j.jtbi.2008.12.018&lt;/a>&lt;/li>
&lt;li>Opris, I., Hampson, R. E., Stanford, T. R., Gerhardt, G. A., &amp;amp; Deadwyler, S. A. (2011). Neural Activity in Frontal Cortical Cell Layers: Evidence for Columnar Sensorimotor Processing. &lt;em>Journal of Cognitive Neuroscience, 23&lt;/em>(6), 1507–1521. &lt;a href="https://doi.org/10.1162/jocn.2010.21534" target="_blank" rel="noopener">https://doi.org/10.1162/jocn.2010.21534&lt;/a>&lt;/li>
&lt;li>Ison, M. J., Mormann, F., Cerf, M., Koch, C., Fried, I., &amp;amp; Quiroga, R. Q. (2011). Selectivity of pyramidal cells and interneurons in the human medial temporal lobe. &lt;em>Journal of Neurophysiology, 106&lt;/em>(4), 1713–1721. &lt;a href="https://doi.org/10.1152/jn.00576.2010" target="_blank" rel="noopener">https://doi.org/10.1152/jn.00576.2010&lt;/a>&lt;/li>
&lt;li>Putrino, D. F., Chen, Z., Ghosh, S., &amp;amp; Brown, E. N. (2011). Motor Cortical Networks for Skilled Movements Have Dynamic Properties That Are Related to Accurate Reaching. &lt;em>Neural Plasticity, 2011&lt;/em>, 1–15. &lt;a href="https://doi.org/10.1155/2011/413543" target="_blank" rel="noopener">https://doi.org/10.1155/2011/413543&lt;/a>&lt;/li>
&lt;/ol></description></item><item><title>Shallow vs. Deep Brain Network Models for Mental Disorder Analysis</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2022-brain-nets/</link><pubDate>Mon, 27 Jun 2022 11:36:49 -0400</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2022-brain-nets/</guid><description>&lt;p>This post was written by Erica Choi, Sally Smith, and Ethan Young and published with minor edits. The team was advised by Professor Carl Yang.
In addition to this post and the &lt;a href="https://www.cs.emory.edu/~jyang71/files/reu2022.pdf" target="_blank" rel="noopener">paper&lt;/a>, the team has also created slides for a &lt;a href="REU_Midterm_Presentation.pdf">midterm presentation&lt;/a>, a &lt;a href="https://youtu.be/DLs1PkO8iJo" target="_blank" rel="noopener">poster blitz video&lt;/a>, and a &lt;a href="REU_Poster.pdf">poster&lt;/a>. This work was accepted at the BrainNN workshop at &lt;a href="Choi-Smith-Young-IEEEBrainNN.pdf">IEEE BigData 2022&lt;/a> and slides are available here.&lt;/p>
&lt;h1 id="comparing-shallow-vs-deep-brain-network-models">Comparing Shallow vs. Deep Brain Network Models&lt;/h1>
&lt;p>Our shallow models use graph kernels to compare structural similarity of brain network data. Plugging those kernels into support vector machines allows us to classify patients&amp;rsquo; brain scans. These kernel methods are called &amp;ldquo;shallow&amp;rdquo; because they do not require many layers of computation, unlike their &amp;ldquo;deep&amp;rdquo; model counterparts. Our deep models are graph neural networks, machine learning (ML) models that can exploit the local information of nodes in graph data to perform classification. In this project, our models are classifying brain scans as either diseased or healthy. We are comparing the two types of models, shallow and deep, to further determine which might be more useful in analyzing neuroimaging data, as well as working on using the models in conjunction with one another to leverage the strengths of both.&lt;/p>
&lt;p>For useful background and definitions refer to &lt;a href="#preliminaries">Preliminaries&lt;/a>.&lt;/p>
&lt;h1 id="datasets">Datasets&lt;/h1>
&lt;p>We are working with 2 datasets, one documenting human immunodeficiency virus (HIV) patients and one documenting bipolar disorder (BP) patients. Each dataset consists of functional magnetic resonance imaging (fMRI) scans, diffusion tensor imaging (DTI) scans, and classification labels in the form of integers, where 1 indicates a healthy patient and -1 indicates an unhealthy patient. Both datasets have been processed for us, as detailed in &lt;a href="https://arxiv.org/abs/2204.07054" target="_blank" rel="noopener">Section 3&lt;/a> of the paper authored by Cui et al.&lt;/p>
&lt;h1 id="problem-formulation">Problem Formulation&lt;/h1>
&lt;p>The DTI and fMRI brain scans of each patient $i$ are represented as weighted adjacency matrices $\mathbf{W}_i \in \mathbb{R}^{M \times M}$. The adjacency matrix is constructed from the brain scan and is a natural way of mathematically representing graph data. Nodes in the brain network represent regions of interest (ROIs), and edge links between nodes indicate the strength of the connection between the two regions. In general, fMRI scans are considered to be more robust than DTI scans; specifically, fMRI scans are less affected by noise caused by data collection. Thus, our experiments prioritize working with the fMRI scans.&lt;/p>
&lt;h2 id="classification-task">Classification Task&lt;/h2>
&lt;p>The standard graph classification task considers the problem of classifying graphs into two or more categories. The goal is to learn a model that maps graphs in the set of graphs $G$ to a set of labels $Y$. In this project, our set of graphs $G$ is the set of brain scans from patients and our set of labels $Y$ consists of two labels: diseased and healthy. The goal of our models is to classify brain scans accurately and improve model interpretability.&lt;/p>
&lt;h2 id="implementation">Implementation&lt;/h2>
&lt;p>For implementation of support vector machines (SVM) with graph kernels, we utilized threshold rounding to remove edge weights and sparsify the adjacency matrices. This means that values in the adjacency matrices were rounded to make the matrices simpler. While this results in information loss, it preserves the overall structure of the adjacency matrices and makes them usable for this particular method. It also makes the computation less expensive. Further manipulation creates a list of graph objects that are compatible with the Python package &lt;a href="https://ysig.github.io/GraKeL/0.1a8/" target="_blank" rel="noopener">GraKel&lt;/a>.&lt;/p>
&lt;p>For implementation of graph convolutional networks (GCNs), we followed &lt;a href="https://github.com/HennyJie/BrainGB" target="_blank" rel="noopener">BrainGB&lt;/a>&amp;rsquo;s code to create a data type that can be used with the Python package &lt;a href="https://pytorch-geometric.readthedocs.io/en/latest/" target="_blank" rel="noopener">PyG&lt;/a>.&lt;/p>
&lt;p>For implementation of kernel graph neural networks (i.e., KerGNN), we followed &lt;a href="https://www.aaai.org/AAAI22Papers/AAAI-6564.FengA.pdf" target="_blank" rel="noopener">KerGNN&lt;/a>&amp;rsquo;s code and implemented threshold rounding to run experiments. The motivation for threshold rounding is the same as for implementing SVM.&lt;/p>
&lt;h1 id="methods">Methods&lt;/h1>
&lt;h2 id="1-graph-kernels">1. Graph Kernels&lt;/h2>
&lt;img src="img/SVC.png" alt="SVC" width="1000"/>
&lt;figcaption align = "center">&lt;b>Fig.1 - Support Vector Machines with Kernels&lt;/b>&lt;/figcaption>
&lt;br/>
&lt;p>We computed three kernels to plug into SVM: Weisfeiler-Lehman (WL), Weisfeiler-Lehman Optimal Assignment (WLOA), and propagation (Prop). The choice of these kernels is motivated by exploiting structural information (i.e., subgraphs) in the brain networks. We tested these graph kernels to find which ones were most effective. On average, the propagation kernel classified HIV best and the WLOA kernel classified bipolar disorder best.&lt;/p>
&lt;h2 id="2-graph-convolutional-networks-gcns">2. Graph Convolutional Networks (GCNs)&lt;/h2>
&lt;img src="img/BrainGB.png" alt="BrainGB" width="1000"/>
&lt;figcaption align = "center">&lt;b>Fig.2 - BrainGB Framework &lt;/b>&lt;/figcaption>
&lt;br/>
&lt;p>The deep model that we experimented with in this project is graph convolutional networks (GCNs). This is a &amp;ldquo;deep&amp;rdquo; model because GCNs are under the umbrella of &amp;ldquo;deep learning&amp;rdquo;. GCNs are modern ML algorithms that pass information through several layers and do more extensive computations than shallow models. Note that our GCNs (and GNNs in general) are shallow in the sense that the models have few layers. Machine learning is still a very active field of research, and recent interest in graph data has led to major strides in graph-based ML.&lt;/p>
&lt;p>We implement message passing GNNs (MPGNN)—a type of GCN—using the BrainGB Python package,
which is built on the Pytorch and Pytorch Geometric libraries. MPGNNs involve what are called message passing schemes to aggregate information from a node&amp;rsquo;s neighbors. Figure 2, adapted from &lt;a href="https://arxiv.org/abs/2204.07054" target="_blank" rel="noopener">Cui et al.&lt;/a>, visualizes the MPGNN architecture.&lt;/p>
&lt;h2 id="3-merging-graph-kernels-and-gnns">3. Merging Graph Kernels and GNNs&lt;/h2>
&lt;p>To leverage the higher order structural information given by graph kernels and local information given by GCNs, we implement GNNs that incorporate various graph kernels (WL, WLOA, etc.) and benchmark their performance on our dataset. The frameworks of particular interest to us are:&lt;/p>
&lt;ul>
&lt;li>the graph convolution layer (GKC) proposed by &lt;a href="https://arxiv.org/abs/2112.07436" target="_blank" rel="noopener">Cosmo et al.&lt;/a>, visualized in Figure 3, and&lt;/li>
&lt;li>the kernel graph neural network (KerGNN) proposed by &lt;a href="https://www.aaai.org/AAAI22Papers/AAAI-6564.FengA.pdf" target="_blank" rel="noopener">Feng et al.&lt;/a>, visualized in Figure 5.&lt;/li>
&lt;/ul>
&lt;img src="img/GKNN.png" alt="Graph Kernel GNN" width="1000"/>
&lt;figcaption align = "center">&lt;b>Fig.3 - GKNN Framework&lt;/b>&lt;/figcaption>
&lt;br/>
&lt;img src="img/KerGNN.png" alt="Graph Kernel GNN" width="1000"/>
&lt;figcaption align = "center">&lt;b>Fig.5 - KerGNN Framework&lt;/b>&lt;/figcaption>
&lt;br/>
&lt;h1 id="results">Results&lt;/h1>
&lt;p>Our most successful model was a graph attention network (GAT) model that used a node concatenation message passing scheme. This model was able to classify HIV patients as healthy or diseased with 81% accuracy on average. In general, our highest performing models were classifying HIV data, particularly using deep models.&lt;/p>
&lt;p>Our highest performing model for bipolar disorder prediction used support vector classifiers (SVCs) with propagation WLOA kernels and had average accuracy of 63%. The differences in performance are minor; furthermore, all kernels&amp;rsquo; mean performance had high standard deviation.&lt;/p>
&lt;p>Our preliminary results from using a combination of kernel methods and GNNs are not outperforming our HIV-GAT(node) model, but we are seeing some improvements in classifying bipolar disorder with the hybrid model, particularly with KerGNN.&lt;/p>
&lt;h1 id="discussion">Discussion&lt;/h1>
&lt;p>In general, we found that our models were better able to classify HIV patients than BP patients. &lt;a href="https://arxiv.org/abs/2204.07054" target="_blank" rel="noopener">Cui et al.&lt;/a> observes that HIV affects both the visual network (VN) and default mode network (DMN), while bipolar disorder mainly affects the bilateral limbic network (BLN). It is possible that HIV was easier to model because it significantly affected multiple networks in the brain, while BP was more elusive with only one major network significantly affected.&lt;/p>
&lt;p>For more details and discussion of our results, see our manuscript (coming soon).&lt;/p>
&lt;h2 id="limitations">Limitations&lt;/h2>
&lt;p>Due to our datasets consisting of less than 100 patients each, our results may not generalize well beyond our specific dataset. If this study were to be replicated, a larger dataset would be ideal, but the expensive nature of brain imaging data and its processing requirements will pose some degree of limitation to any study that uses it. Additionally, brain imaging data is a highly protected data type due to the right to privacy of the patients whose brain scans are used in these experiments. This means that much of the information about the patients is kept private, so it can be challenging to find confounding variables or alternative explanations for statistical results from this data.&lt;/p>
&lt;p>Another notable limitation is that of structure of the brain networks themselves. Specifically, it remains unclear what subgraphs and higher-order information are relevant in classifying brain scans as belonging to diseased or healthy individuals. GNNs also have limitations. For example, GNNs are prone to overfitting, especially with datasets as small as our own. This is an issue that could potentially be alleviated with access to a larger dataset.&lt;/p>
&lt;h2 id="future-work">Future Work&lt;/h2>
&lt;p>There are many avenues with which we may take future research in brain network classification. There are many ways of incorporating graph kernels into GNNs that improve the interpretability of the model, which in turn gives insights into the key underlying structures that help to classify brain networks.&lt;/p>
&lt;h1 id="preliminaries">Preliminaries&lt;/h1>
&lt;p>For technical details of implementing support vector classifiers (SVC) using the Python package &lt;a href="https://scikit-learn.org/stable/" target="_blank" rel="noopener">sklearn&lt;/a>, see this &lt;a href="https://scikit-learn.org/stable/modules/generated/sklearn.svm.SVC.html" target="_blank" rel="noopener">link&lt;/a>.&lt;/p>
&lt;p>For the mathematical theory underlying SVC, see this &lt;a href="https://towardsdatascience.com/support-vector-machine-introduction-to-machine-learning-algorithms-934a444fca47" target="_blank" rel="noopener">blog post&lt;/a>.&lt;/p>
&lt;p>For a survey of graph kernels, see this &lt;a href="https://arxiv.org/abs/1903.11835" target="_blank" rel="noopener">paper&lt;/a>.&lt;/p>
&lt;p>For an introduction to graph neural networks (GNNs), see this &lt;a href="https://distill.pub/2021/gnn-intro/" target="_blank" rel="noopener">blog post&lt;/a>.&lt;/p>
&lt;h1 id="references">References&lt;/h1>
&lt;p>&lt;a href="https://arxiv.org/abs/2204.07054" target="_blank" rel="noopener">BrainGB: A Benchmark for Brain Network Analysis with Graph Neural Networks&lt;/a>&lt;/p>
&lt;p>&lt;a href="https://arxiv.org/abs/2107.05097" target="_blank" rel="noopener">BrainNNExplainer: An Interpretable Graph Neural Network Framework for Brain Network based Disease Analysis&lt;/a>&lt;/p>
&lt;p>&lt;a href="https://dl.acm.org/doi/abs/10.1145/2783258.2783417" target="_blank" rel="noopener">Deep Graph Kernels&lt;/a>&lt;/p>
&lt;p>&lt;a href="https://arxiv.org/abs/2112.07436" target="_blank" rel="noopener">Graph Kernel Neural Networks&lt;/a>&lt;/p>
&lt;p>&lt;a href="https://www.aaai.org/AAAI22Papers/AAAI-6564.FengA.pdf" target="_blank" rel="noopener">KerGNNs: Interpretable Graph Neural Networks with Graph Kernels&lt;/a>&lt;/p></description></item><item><title>From Images to Patient-Specific Models in Cardiology</title><link>http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/</link><pubDate>Tue, 14 Dec 2021 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/</guid><description>&lt;!-- --- -->
&lt;!-- # Emory REU Cardiology Project: -->
&lt;!--Starting new section-->
&lt;!-- --- -->
&lt;p>This post was written by Kai Chang, Allison Dennis, Shannon Lee, Michele Perry, Minxing (Matt) Zhang, and Mohamad Hindawi and published with minor edits. The team was advised by Dr. Alessandro Veneziani.
In addition to this post, the team has also created slides for a &lt;a href="https://docs.google.com/presentation/d/1__H40Xr_KoQaG3Mfhv9aAjzZYjVIvITDuNCAVpsq1v0/edit?usp=sharing" target="_blank" rel="noopener">midterm presentation&lt;/a>, a &lt;a href="https://southalabama.zoom.us/rec/play/nMsrAregiBDRSP8QCj2mDVV7halNAvL0_PvuBcyyf20OraB0BAEtdz7schZwF_Afkmc-ODwH8bNWZ2Q.X_Ym7fzfDZKyswHj?startTime=1627247559000" target="_blank" rel="noopener">poster blitz video&lt;/a>, and a &lt;a href="Emory_Poster.pdf">poster&lt;/a>.&lt;/p>
&lt;h2 id="background-mathematical-modeling-from-the-heart">Background: Mathematical Modeling from the Heart&lt;/h2>
&lt;p>The role of mathematical modeling in clinics is particularly evident in cardiology, as computational mechanics for many historical reasons is a mature field of applied mathematics; on the other hand, many important cardiovascular pathologies have a significant mechanical component, in terms of fluid, structure and their interactions. The clinical impact of mathematical models strongly relies on reconstructing patient geometries to customize and personalize numerical simulations. Advances in medical image processing made over the last two decades have enabled virtual patient-specific models. A key step of the processing pipeline in Cardiology is the extraction of complex vascular geometries like an aortic dissection from medical images (typically, Computed Tomographies, Magnetic Resonance, and Optical Coherence Tomography). Our team evaluated the relation between PDEs and image segmentation/reconstruction through the level set method, and compared this segmentation approach with deep learning ones based on Convolutional Neural Networks.&lt;/p>
&lt;!--Image-->
&lt;p>
&lt;figure id="figure-heart">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="mainImage" srcset="
/site/cmds-reuret/projects/2021-cardio/img/cardio1_hu_371041a6ed02d38d.webp 400w,
/site/cmds-reuret/projects/2021-cardio/img/cardio1_hu_8311f28c587bee29.webp 760w,
/site/cmds-reuret/projects/2021-cardio/img/cardio1_hu_3aca176aed7e4e42.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/img/cardio1_hu_371041a6ed02d38d.webp"
width="384"
height="433"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Heart
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;h2 id="project-overview-comparison-of-segmentation-techniques">Project Overview: Comparison of Segmentation Techniques&lt;/h2>
&lt;p>Imaging has been revolutionizing medical research and clinical practices for decades. One of the components of medical image processing is “segmentation” which has a wide range of applications. Segmentation is the image processing step to identify a region of interest (an artery, a bone, etc.) in an image. We focused on medical images, specifically coronaries based on Optical coherence tomography (OCT). Using OCT images provided by the Emory School of Medicine, we analyzed two segmentation approaches: the Level-Set Method and Machine Learning based on CNN (Convolutional Neural Networks). Our project aimed at segmenting coronaries based on the OCT images using these two methods. We then compared the results of deterministic image-segmentation methods vs deep learning ones.&lt;/p>
&lt;!--Image-->
&lt;p>
&lt;figure id="figure-picture1">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="mainImage" srcset="
/site/cmds-reuret/projects/2021-cardio/img/Picture1_hu_80305eebb338ee61.webp 400w,
/site/cmds-reuret/projects/2021-cardio/img/Picture1_hu_c519f9432e40b8e8.webp 760w,
/site/cmds-reuret/projects/2021-cardio/img/Picture1_hu_98c87b54e64a6fe7.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/img/Picture1_hu_80305eebb338ee61.webp"
width="328"
height="227"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Picture1
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;!--Starting new section-->
&lt;h2 id="materials-and-methods">Materials and Methods&lt;/h2>
&lt;h3 id="level-set-methods-and-vmtk">Level Set Methods and VMTK&lt;/h3>
&lt;p>The Level Set Method utilizes implicit functions to identify the region of interest in an image, where the implicit function is the numerical solution of a Partial Differential Equation (PDE) that is defined on the image that is being segmented. The basic idea of the Level Set is to correlate the velocity to the gray level of the image in such a way that the gray level of the image is driving the evolution of the phi close to the boundary. The Level Set is a very powerful method that extracts the border of a region or image, and it can handle changing topology well. It is also a great tool because it utilizes physical concepts, such as velocity, mean curvature, and elastic energy for image segmentation problems. We used the image segmentation software, Vascular Modeling ToolKit (VMTK), that is based on the Level Set. VMTK is a collection of tools and libraries for image-based modeling of medical images. This segmentation method is model-driven, meaning that the technique is established on physical concepts of the problem.&lt;/p>
&lt;!--Image-->
&lt;p>
&lt;figure id="figure-figure1">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="mainImage" srcset="
/site/cmds-reuret/projects/2021-cardio/img/Figure1_hu_1dddb62d085a8b2d.webp 400w,
/site/cmds-reuret/projects/2021-cardio/img/Figure1_hu_1758ac0a7e136851.webp 760w,
/site/cmds-reuret/projects/2021-cardio/img/Figure1_hu_a420b158ec85bd35.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/img/Figure1_hu_1dddb62d085a8b2d.webp"
width="288"
height="209"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Figure1
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;h3 id="fenics-and-matlab">FEniCS and MATLAB&lt;/h3>
&lt;p>The FEniCS Project is a research and software project aimed at creating mathematical methods and software for automated computational mathematical modeling. As the implicit function, the numerical solution of a PDE, is involved in our project, we tried to use FEniCS in Python for solving PDEs using finite element methods. For MATLAB, our group used Image Segmenter App under Image Processing Toolbox and applied Thresholding, Active Contours, Graph Cut, Auto Cluster, etc. to segment 2-D images. For 3-D volumetric images, we used Volume Segmenter to create and refine binary or semantic segmentation masks to segment the images by means of automated, semi-automated, and manual techniques.&lt;/p>
&lt;!--Image-->
&lt;p>
&lt;figure id="figure-figure2">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="mainImage" srcset="
/site/cmds-reuret/projects/2021-cardio/img/Figure2_hu_f2f1fb67473eece5.webp 400w,
/site/cmds-reuret/projects/2021-cardio/img/Figure2_hu_b75552a9164181c2.webp 760w,
/site/cmds-reuret/projects/2021-cardio/img/Figure2_hu_9f6f69b4cc60f974.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/img/Figure2_hu_f2f1fb67473eece5.webp"
width="216"
height="184"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Figure2
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;h3 id="convolutional-neural-networks-cnns">Convolutional Neural Networks (CNNs)&lt;/h3>
&lt;p>Convolutional Neural Networks are a deep learning algorithm for image classification. The CNN&amp;rsquo;s convolutional layer parameters comprise of filters, where the values of the filters are learned during the training phase. The layers are for feature learning and classification, specifically for classifying the pixels in an image with respect to a background or vessel. We used the image processing software, PyTorch. The deep learning (DL) based method involves using training data from a database of images to train the algorithm in PyTorch.&lt;/p>
&lt;!--Image-->
&lt;p>
&lt;figure id="figure-figure7">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="mainImage" srcset="
/site/cmds-reuret/projects/2021-cardio/img/Figure7_hu_17e1a8d4cf35c440.webp 400w,
/site/cmds-reuret/projects/2021-cardio/img/Figure7_hu_968ca732578b34b.webp 760w,
/site/cmds-reuret/projects/2021-cardio/img/Figure7_hu_64899c76a3d7761f.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/img/Figure7_hu_17e1a8d4cf35c440.webp"
width="288"
height="313"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Figure7
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;h2 id="our-activities--experiences">Our Activities &amp;amp; Experiences&lt;/h2>
&lt;h3 id="weeks-1-and-2">Weeks 1 and 2&lt;/h3>
&lt;p>Using CNNs, we generated OCT segmentation maps in PyTorch on the original data provided by Dr. Molony. The same data were used in VMTK slice by slice.&lt;/p>
&lt;h3 id="week-3">Week 3&lt;/h3>
&lt;p>We added 15 new images to our OCT data set in PyTorch. With more training data, the loss value went down, and the accuracy increased. We worked to find the most accurate segmentation process in VMTK to extract the level set for the boundary of the coronary. We cleaned a sample image using GIMP. To create a 3D stack of images, each slice was replicated 30 times. After extracting the surface in VMTK, ParaView extracted the outline of the region. We were then in a good position to start comparing the contours from the two methods.&lt;/p>
&lt;h3 id="research-week-4">Research Week 4:&lt;/h3>
&lt;p>We explored different metrics for comparing the results visually and numerically. One option is the Jaccard Index, which quantifies the percent overlap between the segmented images. Another option was to extract the VMTK and PyTorch contour lines and boundary points in ParaView and then use Python or Matlab code to plot the x and y coordinates of the results as overlapping figures. Then, we could use integration to calculate the difference between the two curves (L^2 metrics) .&lt;/p>
&lt;!--Image-->
&lt;p>
&lt;figure id="figure-illustration2">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="mainImage" srcset="
/site/cmds-reuret/projects/2021-cardio/img/illustration2_hu_9ebcd96f90cfce72.webp 400w,
/site/cmds-reuret/projects/2021-cardio/img/illustration2_hu_41eabc72ab794c91.webp 760w,
/site/cmds-reuret/projects/2021-cardio/img/illustration2_hu_b5041fb655f37508.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/img/illustration2_hu_9ebcd96f90cfce72.webp"
width="432"
height="532"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Illustration2
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;!--Starting new section-->
&lt;hr>
&lt;h2 id="comparison">Comparison&lt;/h2>
&lt;h3 id="which-method-was-faster">Which method was faster?&lt;/h3>
&lt;p>The deep learning method was much faster than VMTK, but it still required a waiting period to complete the epochs. This might be because our training data and development data was relatively small. We came to recognize that a limitation of the deep learning approach was that labeled and large data sets are required for preventing CNN overfitting and increasing accuracy.&lt;/p>
&lt;h3 id="which-method-required-more-human-intervention">Which method required more human intervention?&lt;/h3>
&lt;p>VMTK segmentation required far more human intervention than the DL method. With VMTK, we had to initialize the image by selecting an initialization type, identifying a lower and upper threshold of pixel values, and placing seeds that identified the region we planned to segment. We also selected the numerical values for the level set conditions, including number of iterations, propagation, curvature, and advection. We actively supervised the segmentation and chose whether to accept or reject the results. On the other hand, the DL method required much less decision making. With the DL algorithm, we empirically set a few parameters associated with training the CNN, such as training rate and number of epochs. Then, the DL algorithm trained itself and segmented the OCT images without user participation.&lt;/p>
&lt;h3 id="other-advantages-and-limitations">Other advantages and limitations?&lt;/h3>
&lt;p>PyTorch is better at handling noise suppression than VMTK. On the other hand, VMTK handles the data normalization, contrast enhancement, and conversion of color images to grayscale better.&lt;/p>
&lt;!--Image-->
&lt;p>
&lt;figure id="figure-illustration">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="mainImage" srcset="
/site/cmds-reuret/projects/2021-cardio/img/illustration_hu_c4c56fdc8f74fc1.webp 400w,
/site/cmds-reuret/projects/2021-cardio/img/illustration_hu_7f661e86cdb1ab2c.webp 760w,
/site/cmds-reuret/projects/2021-cardio/img/illustration_hu_b9410da47979a361.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/projects/2021-cardio/img/illustration_hu_c4c56fdc8f74fc1.webp"
width="648"
height="288"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
Illustration
&lt;/figcaption>&lt;/figure>
&lt;/p>
&lt;h2 id="our-results--conclusions">Our Results &amp;amp; Conclusions&lt;/h2>
&lt;p>After evaluating the benefits and limitations of both methods, we have come to the conclusion that instead of preferring one over the other, we can combine the two for better results. We could train a neural network with an optimal parameter selection to use in VMTK segmentation for setting the parameters. We can have a CNN trained on the OCT data to tell us what the correct parameter set is. This could prove to be useful, especially since we encountered a set-back in determining the initialization type, thresholds, and parameters to set for the best segmentation result in VMTK. We should not have a competition over the two techniques, rather, we should combine the strengths and benefits of both model-driven and data-driven approaches.&lt;/p>
&lt;p>Future works should focus on improved segmentation using unsupervised Deep Learning where the machine uses image-derived features, or supervised learning that requires Gold Standard (GS) segmentation to train it. The deep learning-based algorithm demonstrated high accuracy based on Jaccard Index. We should look into Edge-based deformable models, and approaches using blood vessel tracking algorithms and seeding points to find the minimum path according to image-derived metrics.&lt;/p>
&lt;!--Insert Image Here-->
&lt;!--Starting new section-->
&lt;!--Starting new section-->
&lt;h2 id="more-about-the-team">More About The Team&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Allison Dennis&lt;/strong> is a rising junior at Texas A&amp;amp;M University majoring in Applied Mathematics with a minor in Cybersecurity. Her interests are scientific computing and linear algebra, and her dream career is to be a Calculus 1 and 2 professor. Besides doing math, she loves to read, hike, watch Gilmore Girls, and spend time with her family and friends.&lt;/li>
&lt;li>&lt;strong>Dr. Mohamad Hindawi&lt;/strong> teaches AP Calc in Tucker High school. He is certified to teach AP Physics, AP Chem, AP Stat, and AP Comp Sc. His interest is in Mathematical Modeling and Differential Eq specially Navier-Stokes Equation. Outside academia, He enjoys swimming long distances, deep sea water fishing in Alaska for Halibut and King Salmon.&lt;/li>
&lt;li>&lt;strong>Shannon Lee&lt;/strong> is a rising junior at Southern Methodist University majoring in accounting, applied mathematics, and statistics. Her interests are learning new computational math and statistical modeling techniques in all areas. Outside of school, she enjoys playing tennis, traveling, and spending time with friends and family.&lt;/li>
&lt;li>&lt;strong>Michele Perry&lt;/strong> is a rising senior at University of South Alabama majoring in Math/Statistics and minoring in music. Her dream career is to utilize math in astronomical research at NASA or a university. Outside of academics, she spends her time practicing for band, reading, traveling, and hanging out with her cat and dogs.&lt;/li>
&lt;li>&lt;strong>Minxing(Matt) Zhang&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Kai Chang&lt;/strong>&lt;/li>
&lt;/ul>
&lt;!--Starting new section-->
&lt;h1 id="references">References&lt;/h1>
&lt;ol>
&lt;li>&lt;a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3785070/" target="_blank" rel="noopener">https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3785070/&lt;/a>&lt;/li>
&lt;li>Moccia, Sara, et al. &amp;ldquo;Blood vessel segmentation algorithms—review of methods, datasets and evaluation metrics.&amp;rdquo; Computer methods and programs in biomedicine, 158 (2018): 71-91.&lt;/li>
&lt;li>&lt;a href="https://fenicsproject.org" target="_blank" rel="noopener">https://fenicsproject.org&lt;/a>&lt;/li>
&lt;/ol></description></item><item><title>Iterative Methods at Lower Precision</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/chen-et-al-2023/</link><pubDate>Fri, 01 Aug 2031 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/chen-et-al-2023/</guid><description>&lt;p>Since numbers in the computer are represented with a fixed number of bits, loss of accuracy during calculation is unavoidable. At high precision where more bits (e.g. 64) are allocated to each number, round-off errors are typically small. On the other hand, calculating at lower precision, such as half (16 bits), has the advantage of being much faster. This research focuses on experimenting with arithmetic at different precision levels for large-scale inverse problems, which are represented by linear systems with ill-conditioned matrices. We modified the Conjugate Gradient Method for Least Squares (CGLS) and the Chebyshev Semi-Iterative Method (CS) with Tikhonov regularization to do arithmetic at lower precision using the MATLAB chop function, and we ran experiments on applications from image processing and compared their performance at different precision levels. We concluded that CGLS is a more stable algorithm, but overflows easily due to the computation of inner products, while CS is less likely to overflow but it has more erratic convergence behavior. When the noise level is high, CS outperforms CGLS by being able to run more iterations before overflow occurs; when the noise level is close to zero, CS appears to be more susceptible to accumulation of round-off errors.&lt;/p></description></item><item><title>Call for Summer 2026 Applications</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2026-cfa/</link><pubDate>Fri, 19 Dec 2025 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2026-cfa/</guid><description>&lt;p>We are pleased to announce that our applications website for &lt;a href="https://etap.nsf.gov/award/7517/opportunity/11727" target="_blank" rel="noopener">REU applicants&lt;/a> is now live. We will start reviewing applications on February 1 and all applications received by March 1 will receive full consideration.&lt;/p>
&lt;p>Our theme will be &lt;em>Computational Science and AI for New Mathematics&lt;/em>. This year&amp;rsquo;s projects span both pure and applied mathematics, exploring how modern AI can push the boundaries of mathematical discovery. For more information about the application process, see &lt;a href="../../apply/">this page&lt;/a> and for a list of projects see the &lt;a href="../../summer2026">Summer 2026 tab&lt;/a>.&lt;/p>
&lt;p>For questions about the program, please contact &lt;strong>Dr. Levon Nurbekyan&lt;/strong> at &lt;a href="mailto:levon.nurbekyan@emory.edu">levon.nurbekyan@emory.edu&lt;/a>.&lt;/p>
&lt;h2 id="2026-project-mentors">2026 Project Mentors&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Dr. Levon Nurbekyan&lt;/strong> - &lt;a href="../../projects/2026-faber-krahn/">AI-Assisted Exploration of the Polygonal Faber-Krahn Inequality&lt;/a>&lt;/li>
&lt;li>&lt;strong>Dr. Marco Tezzele&lt;/strong> and &lt;strong>Dr. Jimena Martin Tempestti&lt;/strong> - &lt;a href="../../projects/2026-digital-twins/">Generative AI for Structure Discovery in Cardiovascular Digital Twins&lt;/a>&lt;/li>
&lt;li>&lt;strong>Dr. Deependra Singh&lt;/strong> - &lt;a href="../../projects/2026-algebra-nt/">AI-Assisted Exploration in Algebra and Number Theory&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Applications of Automatic Differentiation in Image Registration</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/watson-et-al-2024/</link><pubDate>Mon, 28 Jul 2025 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/watson-et-al-2024/</guid><description>&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/8O9S8zm2N-E" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen>&lt;/iframe>
&lt;p>We demonstrate that automatic differentiation, which has become commonly available in machine learning frameworks, is an efficient way to explore ideas that lead to algorithmic improvement in multi-scale affine image registration and affine super-resolution problems. In our first experiment on multi-scale registration, we implement an ODE predictor-corrector method involving a derivative with respect to the scale parameter and the Hessian of an image registration objective function, both of which would be difficult to compute without AD. Our findings indicate that exact Hessians are necessary for the method to provide any benefits over a traditional multi-scale method; a Gauss-Newton Hessian approximation fails to provide such benefits. In our second experiment, we implement a variable projected Gauss-Newton method for super-resolution and use AD to differentiate through the iteratively computed projection, a method previously unaddressed in the literature. We show that Jacobians obtained without differentiating through the projection are poor approximations to the true Jacobians of the variable projected forward map and explore the performance of some other approximations. By addressing these problems, this work contributes to the application of AD in image registration and sets a precedent for further use of machine learning tools in this field.&lt;/p></description></item><item><title>Call for Summer 2025 Applications</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2025-cfa/</link><pubDate>Thu, 19 Dec 2024 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2025-cfa/</guid><description>&lt;p>We are pleased to announce that our applications website for &lt;a href="https://etap.nsf.gov/award/7517/opportunity/10142" target="_blank" rel="noopener">REU applicants&lt;/a> is now live. As we did last year, we are using NSF&amp;rsquo;s new ETAP system for the first time, which should make it easy for applicants to apply. We will start reviewing applications on February 1 and all applications received by March 1 will receive full consideration.&lt;/p>
&lt;p>Our theme will be &lt;em>Models meet Data&lt;/em>. For more information about the application process, see &lt;a href="../../apply/">this page&lt;/a> and for a list of projects see the &lt;a href="../../summer2025">Summer 2025 tab&lt;/a>.&lt;/p></description></item><item><title>Highlights from Our 2024 Summer REU Program</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2024-summary/</link><pubDate>Thu, 07 Nov 2024 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2024-summary/</guid><description>&lt;p>This summer, we welcomed 15 talented students from across the country on Emory&amp;rsquo;s beatiful campus. Under the umbrella theme Learning from Images, the students worked in five teams, each diving into projects aimed at advancing our understanding and ability to learn from imaging data. This program, held over eight weeks on campus, provided the young researchers with an opportunity to engage in cutting-edge computational mathematics and machine learning.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Group picture of the 2024 mentors and students" srcset="
/site/cmds-reuret/news/2024-summary/DSCF1058_hu_5b60b25127d1b67.webp 400w,
/site/cmds-reuret/news/2024-summary/DSCF1058_hu_cc7a1b27d6350e7b.webp 760w,
/site/cmds-reuret/news/2024-summary/DSCF1058_hu_d7f426e979aaa70f.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/news/2024-summary/DSCF1058_hu_5b60b25127d1b67.webp"
width="760"
height="418"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>The Learning from Images theme tackled a timely challenge: how to make sense of the overwhelming amount of imaging data generated daily. Computational algorithms can unlock insights from this data, improving our ability to detect new patterns and enhance image quality, especially in areas critical to healthcare. Each project addressed a unique aspect of this challenge—from improving image distribution models and optimizing measurement designs to refining algorithms for reconstructing image sequences. You can read more about the projects &lt;a href="../../summer2024/">here&lt;/a>.&lt;/p>
&lt;p>The students worked with their mentors to explore and develop advanced imaging techniques that leverage machine learning, numerical linear algebra, optimization, and differential equations. Weekly project meetings, a seminar series, and various social events fostered an engaging learning environment. Students had the chance to present their progress in midterm presentations, culminating in a highly anticipated final poster session.&lt;/p>
&lt;p>The poster session marked the program’s grand finale, drawing a large crowd of faculty, students, and postdocs. Each team showcased their findings and shared insights from their experiments, highlighting the strengths and potential areas of improvement for their algorithms. By a narrow margin, this year’s poster award went to the team of Malia Walewski, Rishi Leburu, Claire Gan, and Callihan Bertley for their project on &lt;a href="../../projects/2024-vae/">Improving VAEs with Normalizing Flows&lt;/a>. The picture below shows Malia and Rishi presenting their work.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="Malia and Rishi presenting their work" srcset="
/site/cmds-reuret/news/2024-summary/DSCF1043_hu_1b6d08fca9d78e22.webp 400w,
/site/cmds-reuret/news/2024-summary/DSCF1043_hu_995e88e22fabceed.webp 760w,
/site/cmds-reuret/news/2024-summary/DSCF1043_hu_c149dec8f78bbc92.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/news/2024-summary/DSCF1043_hu_1b6d08fca9d78e22.webp"
width="760"
height="716"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>All teams created a poster blitz video and wrote blog posts about their work, which are now live on the &lt;a href="../../summer2024/">2024 tab&lt;/a>. These additional resources capture the excitement and innovation of the program and give a closer look into the students’ experiences.&lt;/p>
&lt;p>We’re proud of our 2024 REU cohort and look forward to following their careers in the years to come!&lt;/p></description></item><item><title>Training Implicit Networks for Image Deblurring using Jacobian-Free Backpropagation</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/liu-et-al-2022/</link><pubDate>Thu, 01 Feb 2024 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/liu-et-al-2022/</guid><description>&lt;p>Recent efforts in applying implicit networks to solve inverse problems in imaging have achieved competitive or even superior results when compared to feedforward networks. These implicit networks only require constant memory during backpropagation, regardless of the number of layers. However, they are not necessarily easy to train. Gradient calculations are computationally expensive because they require backpropagating through a fixed point. In particular, this process requires solving a large linear system whose size is determined by the number of features in the fixed point iteration. This paper explores a recently proposed method, Jacobian-free Backpropagation (JFB), a backpropagation scheme that circumvents such calculation, in the context of image deblurring problems. Our results show that JFB is comparable against fine-tuned optimization schemes, state-of-the-art (SOTA) feedforward networks, and existing implicit networks at a reduced computational cost.&lt;/p></description></item><item><title>Call for Summer 2024 Applications</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2024-cfa/</link><pubDate>Thu, 21 Dec 2023 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2024-cfa/</guid><description>&lt;p>We are pleased to announce that our applications website for &lt;a href="https://etap.nsf.gov/award/5215/opportunity/6964" target="_blank" rel="noopener">REU applicants&lt;/a> is now live. This year, we are using NSF&amp;rsquo;s new ETAP system for the first time, which should make it easier for applicants to apply. We will start reviewing applications on February 1 and all applications received by March 1 will receive full consideration.&lt;/p>
&lt;p>For more information about the application process, see &lt;a href="../../apply/">this page&lt;/a> and stay tuned for a list of projects to be posted a the &lt;a href="../../summer2024">Summer 2024 tab&lt;/a>.&lt;/p>
&lt;h1 id="learning-from-images">Learning from Images&lt;/h1>
&lt;p>The amount of imaging data generated every day exceeds human imagination. With their ability to statistically analyze such large datasets, computational algorithms can enhance our ability to discover new patterns and improve imaging data quality for critical healthcare applications and beyond.&lt;/p>
&lt;p>This theme&amp;rsquo;s projects provide new mathematical insights and algorithms enabling learning from image data. The goals of the individual projects include improving algorithms for learning the distribution of image data, optimizing the measurement design to improve image quality, developing efficient algorithms for reconstructing image sequences, and generalizing machine learning techniques to learn transformations between images.&lt;/p>
&lt;p>The projects build upon and advance state-of-the-art techniques from machine learning, numerical linear algebra, and differential equations. Students will learn about these techniques and be trained to combine them in new ways to build effective algorithms. Their analysis and experiments will provide new insights into the strengths and weaknesses of their approaches.&lt;/p></description></item><item><title>MCRAGE: Synthetic Healthcare Data for Fairness</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/behal-et-al-2023/</link><pubDate>Fri, 27 Oct 2023 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/behal-et-al-2023/</guid><description>&lt;iframe width="560" height="315" src="https://www.youtube.com/watch?v=eyijGEz9CZg" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen>&lt;/iframe>
&lt;p>In the field of healthcare, electronic health records (EHR) serve as crucial training data for developing machine learning models for diagnosis, treatment, and the management of healthcare resources. However, medical datasets are often imbalanced in terms of sensitive attributes such as race/ethnicity, gender, and age. Machine learning models trained on class-imbalanced EHR datasets perform significantly worse in deployment for individuals of the minority classes compared to samples from majority classes, which may lead to inequitable healthcare outcomes for minority groups. To address this challenge, we propose Minority Class Rebalancing through Augmentation by Generative modeling (MCRAGE), a novel approach to augment imbalanced datasets using samples generated by a deep generative model. The MCRAGE process involves training a Conditional Denoising Diffusion Probabilistic Model (CDDPM) capable of generating high-quality synthetic EHR samples from underrepresented classes. We use this synthetic data to augment the existing imbalanced dataset, thereby achieving a more balanced distribution across all classes, which can be used to train an unbiased machine learning model. We measure the performance of MCRAGE versus alternative approaches using Accuracy, F1 score and AUROC. We provide theoretical justification for our method in terms of recent convergence results for DDPMs with minimal assumptions.&lt;/p></description></item><item><title>Emory Welcomes 2023 REU/RET Participants on Campus</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2023-kickoff/</link><pubDate>Tue, 13 Jun 2023 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2023-kickoff/</guid><description>&lt;p>This is an exciting week for us as we finally entered the core phase of our 2023 REU/RET.
We are delighted to have a diverse and &lt;a href="../../people/">talented group of participants&lt;/a> joining us from different states and institutions.
They will spend the next six weeks on Emory’s campus and will have access to state-of-the-art facilities, resources, and mentorship.
We look forward to getting to know them and supporting them throughout this journey.&lt;/p>
&lt;p>The theme of this year’s program is data for social justice and our five teams will explore various topics related to this theme from a mathematical perspective.
Our &lt;a href="../../summer2023/">projects&lt;/a> span topics motivated by health disparities, environmental justice, and fairness in deep learning.
The participants will study and develop mathematical algorithms and computational tools to analyze and solve open challenges in these areas.
They will have plenty of opportunities to meet with their mentors and peers who will provide feedback and guidance.&lt;/p>
&lt;p>In the next six weeks, our participants will not only work on their research projects, but also engage in a variety of activities that will enhance their skills and knowledge in data science.
Our weekly seminar and ad-hoc lectures will train participants in most aspects of mathematical research in this area and give advice on career development.&lt;/p>
&lt;p>We are excited to kick off this program and look forward to a productive and fun summer with our participants!&lt;/p></description></item><item><title>Ensemble Kalman Filtering for Glacier Modeling</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/corcoran-et-al-2022/</link><pubDate>Thu, 01 Dec 2022 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/corcoran-et-al-2022/</guid><description>&lt;p>Working with a two-stage ice sheet model, we explore how statistical data assimilation methods can be used to improve predictions of glacier melt and relatedly, sea level rise. We find that the EnKF improves model runs initialized using incorrect initial conditions or parameters, providing us with better models of future glacier melt. We explore the necessary number of observations needed to produce an accurate model run. Further, we determine that the deviations from the truth in output that stem from having few data points in the pre-satellite era can be corrected with modern observation data. Finally, using data derived from our improved model we calculate sea level rise and model storm surges to understand the affect caused by sea level rise.&lt;/p></description></item><item><title>Participants present their work at the REU/RET Poster Session</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2022-posteraward/</link><pubDate>Sun, 26 Jun 2022 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2022-posteraward/</guid><description>&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="&amp;ldquo;Participants present their work at the REU/RET Poster session.&amp;rdquo;" srcset="
/site/cmds-reuret/news/2022-posteraward/PosterSession_hu_91335fac56652ac1.webp 400w,
/site/cmds-reuret/news/2022-posteraward/PosterSession_hu_ab92d7dd54ded8cc.webp 760w,
/site/cmds-reuret/news/2022-posteraward/PosterSession_hu_7b9d40a8f090e2be.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/news/2022-posteraward/PosterSession_hu_91335fac56652ac1.webp"
width="760"
height="438"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>In their last week on campus, our REU/RET teams presented their work in our annual poster session.
The event was held in our beautiful atrium and was a magnet for many students, postdocs, and faculty from the Departments of Mathematics and Computer Science.&lt;/p>
&lt;p>In preparation for the event and to advertise their work, the teams made two-minute Poster Blitz videos.
Since the videos were full of creativity, we played them during the event using our smart screen. The videos are now hosted at:&lt;/p>
&lt;ul>
&lt;li>&lt;a href="https://www.youtube.com/watch?v=2hDyfdaM5Es" target="_blank" rel="noopener">Hamiltonian Neural Networks&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://www.youtube.com/watch?v=bGeOZ9G6IOc" target="_blank" rel="noopener">Storm Surge modeling&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://youtu.be/rD83floj3Jg" target="_blank" rel="noopener">Mixed Precision Linear Algebra&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://youtu.be/2uBVgNFRpqI" target="_blank" rel="noopener">Modeling Neural Firing&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://youtu.be/oIwL3E2yULg" target="_blank" rel="noopener">Fast Training of Implicit Neural Nets&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://youtu.be/DLs1PkO8iJo" target="_blank" rel="noopener">Deep Learning for Mental Disorder Analysis&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://youtu.be/i9g6mRNJEHA" target="_blank" rel="noopener">Reinforcement Learning vs Optimal Control&lt;/a>&lt;/li>
&lt;/ul>
&lt;p>During the event, more than a dozen judges had the difficult task of scoring the quality of the posters, how well the teams presented and explained their work, and the results of the research itself, to determine the winning team.
All teams scored very well on and with the smallest of margins, the poster on &lt;strong>reinforcement learning vs. optimal control&lt;/strong> won the award.
The team consists of &lt;a href="https://www.linkedin.com/in/arjunso/" target="_blank" rel="noopener">Arjun Sethi-Olowin&lt;/a> (a rising senior at Rice University), &lt;a href="https://dewanchowdhury.github.io/" target="_blank" rel="noopener">Dewan Chowdhury&lt;/a> (a rising junior at Rutgers University), and &lt;a href="https://www.linkedin.com/in/jacob-mantooth-7b262321b" target="_blank" rel="noopener">Jacob Mantooth&lt;/a> (a rising senior at East Central University). They each received a $50 gift card from a major online retailer.&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="&amp;ldquo;Arjun Sethi-Olowin, Jacob Mantooh, and Dewan Chowdhury receive the best poster award.&amp;rdquo;" srcset="
/site/cmds-reuret/news/2022-posteraward/IMG_0645_hu_31501ae8f5a1fc9e.webp 400w,
/site/cmds-reuret/news/2022-posteraward/IMG_0645_hu_90e39291d967ea46.webp 760w,
/site/cmds-reuret/news/2022-posteraward/IMG_0645_hu_dbc03cb52a151065.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/news/2022-posteraward/IMG_0645_hu_31501ae8f5a1fc9e.webp"
width="760"
height="389"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>But no one left empty-handed! Our official REU/RET t-shirts, designed by our outreach committee, arrived just in time. Also, we held a raffle with Emory and SIAM merchandise. Here, Jim Nagy presents one of the main prizes, a pair of SIAM socks, going to Logan Knudsen:&lt;/p>
&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="&amp;ldquo;Logan Knudsen receives SIAM socks.&amp;rdquo;" srcset="
/site/cmds-reuret/news/2022-posteraward/IMG_0642_hu_c3528edfb9ee507.webp 400w,
/site/cmds-reuret/news/2022-posteraward/IMG_0642_hu_b4fa8edbd7456a2c.webp 760w,
/site/cmds-reuret/news/2022-posteraward/IMG_0642_hu_ec05d0d5bceda190.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/news/2022-posteraward/IMG_0642_hu_c3528edfb9ee507.webp"
width="760"
height="452"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>We thank the Department of Mathematics for making the poster session a huge hit and providing the prizes.&lt;/p></description></item><item><title>Emory Welcomes 2022 REU/RET Participants on Campus</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2022-kickoff/</link><pubDate>Mon, 13 Jun 2022 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2022-kickoff/</guid><description>&lt;p>
&lt;figure >
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="&amp;ldquo;2022 REU/RET participants arrive on Emory Campus&amp;rdquo;" srcset="
/site/cmds-reuret/news/2022-kickoff/IMG_5393_hu_6fe6db2f62806be0.webp 400w,
/site/cmds-reuret/news/2022-kickoff/IMG_5393_hu_ffb7b6565981c629.webp 760w,
/site/cmds-reuret/news/2022-kickoff/IMG_5393_hu_94237c8eb4e9b7cc.webp 1200w"
src="http://www.math.emory.edu/site/cmds-reuret/news/2022-kickoff/IMG_5393_hu_6fe6db2f62806be0.webp"
width="760"
height="444"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;/p>
&lt;p>Today marks the first day of our core REU/RET phase 2022, and we are delighted to welcome our participants from across the country to our beautiful campus. This year on our site, 22 undergraduate researchers and four teachers will work on eight projects under the theme &lt;a href="../../summer2022">&lt;em>Models meet Data&lt;/em>&lt;/a>. All projects will combine tools from computational mathematics (modeling, differential equations, optimization, linear algebra) with statistical and data science techniques (deep learning, reinforcement learning, data assimilation) to tackle challenging research questions.&lt;/p>
&lt;p>Over the next six weeks, the groups will have several milestones which give opportunities to practice most modes of scientific communication. In their research contracts, the participants will organize their collaboration and develop a timeline for the project. Their midterm presentations give an excellent opportunity to expose synergies and ideas across the teams. The participants will also create a blog post about their work, and there will be a poster session toward the end of the program. Throughout the six weeks, the teams will work on a manuscript that describes their project.&lt;/p>
&lt;p>In addition to working on this packed research schedule, we will have many events to foster learning (e.g., ad-hoc lectures, a weekly seminar), networking (e.g., excursions and other social events organized by the participants), and professional development (e.g., scientific writing workshops, career panels).&lt;/p>
&lt;p>Besides the NSF support, we are also grateful for the support from Emory&amp;rsquo;s &lt;a href="http://math.emory.edu" target="_blank" rel="noopener">Department of Mathematics&lt;/a>, which provides the top-notch facilities to run this program smoothly and funds to host events such as this morning&amp;rsquo;s breakfast. In addition to one co-working area, a computer pool, and two seminar rooms dedicated to our site, we&amp;rsquo;ll also try out our department&amp;rsquo;s computing facilities and new digital whiteboard.&lt;/p></description></item><item><title>REU Participants Receive National Fellowships</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2022-congrats-to-fellows/</link><pubDate>Fri, 20 May 2022 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2022-congrats-to-fellows/</guid><description>&lt;p>We are excited to congratulate Manuel Santana and Katie Keegan, who participated in our program in the summer of 2021, on winning nationally-competitive scholarships that will accelerate enhance their graduate school career.&lt;/p>
&lt;p>Manuel Santana, a &lt;a href="https://www.usu.edu/today/story/gold-standard-two-aggie-mathematicians-are-2021-goldwater-scholars" target="_blank" rel="noopener">2021 Goldwater Scholar&lt;/a>, who graduated with a BA in Computational Mathematics from Utah State University was awarded a prestigious scholarship through the &lt;a href="https://www.nsfgrfp.org/" target="_blank" rel="noopener">NSF&amp;rsquo;s Graduate Research Fellowship Program (GRFP)&lt;/a>. Besides his research on computational methods that &lt;a href="../../projects/2021-tomography">accelerate the reconstruction of computed tomography&lt;/a> on our site, he also worked on &lt;a href="https://arxiv.org/abs/2012.10591" target="_blank" rel="noopener">combinatorics&lt;/a>, and shared many &lt;a href="https://www.numerade.com/educators/profile/manuel-s-36048/?page=22" target="_blank" rel="noopener">tutorials for solving algebra problems&lt;/a>. He will join Cal Tech for his PhD.&lt;/p>
&lt;p>&lt;a href="https://katiekeegan.org/" target="_blank" rel="noopener">Katie Keegan&lt;/a> was offered not only an NSF GRFP but also a [Computational Science Graduate Fellowship by the US Department of Energy]0(&lt;a href="https://www.krellinst.org/csgf/%29" target="_blank" rel="noopener">https://www.krellinst.org/csgf/)&lt;/a>, which she accepted. Katie completed her BS in Applied Mathematics at Mary Baldwin College this spring and participated in two REUs. In the summer of 2020, she was part of ICERM&amp;rsquo;s REU and published a paper on a &lt;a href="https://www.siam.org/Portals/0/Documents/S141166PDF.pdf?ver=2021-09-23-070730-093" target="_blank" rel="noopener">modified watermarking scheme based on the SVD&lt;/a>, which was later also featured in &lt;a href="https://sinews.siam.org/Details-Page/a-modified-watermarking-scheme-based-on-the-singular-value-decomposition" target="_blank" rel="noopener">SIAM News&lt;/a>. On our site she developed a &lt;a href="../../projects/2021-tensor">tensor SVD-based classification algorithm for fMRI data&lt;/a>. We are extremely delighted that Katie will join our Computational Mathematics PhD program starting this fall.&lt;/p></description></item><item><title>2022 Projects and Mentors Announced</title><link>http://www.math.emory.edu/site/cmds-reuret/news/2022-projects/</link><pubDate>Thu, 24 Feb 2022 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/news/2022-projects/</guid><description>&lt;p>With the application deadline for the summer 2022 program approaching, we&amp;rsquo;re delighted to announce the project topics and mentors. This year, we will have six projects centered on the theme models meet data. Applications include storm surge modeling and neuroscience and mathematical techniques cover a variety of topics ranging from data assimilation to deep neural networks and reinforcement learning. The teams will be mentored by faculty from Emory&amp;rsquo;s department of Mathematics and Computer Science and two former faculty and students who will join us as guest mentors. A full list of topics can be found on the &lt;a href="../../summer2022/">Summer 2022 page&lt;/a>. Applications for the &lt;a href="https://www.mathprograms.org/db/programs/1215" target="_blank" rel="noopener">REU &lt;/a> and &lt;a href="https://www.mathprograms.org/db/programs/1214" target="_blank" rel="noopener">RET &lt;/a> are due on March 1.&lt;/p></description></item><item><title>Comparing Shallow and Deep Graph Models for Brain Network Analysis</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/choi-et-al-2022/</link><pubDate>Tue, 04 Jan 2022 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/choi-et-al-2022/</guid><description>&lt;p>It is well-established that graph neural networks (GNNs) can effectively model networked data in a variety of fields. However, whether GNNs can outperform traditional shallow graph classification models such as graph kernels for brain network analysis remains unclear. To this end, we analyze different approaches for modeling brain networks, including graph kernel based SVM, basic GNNs and kernelized GNNs. These models are designed to aid in the analysis of diseases and mental disorders such as bipolar disorder, human immunodeficiency virus (HIV), post-traumatic stress disorder (PTSD), and depression. In particular, we conduct experiments with three methods: kernelized support vector machines (SVM), message passing graph neural networks (MPGNNs), and kernel graph neural networks (KerGNN). We conclude that 1) deep models (GNNs) generally outperform shallow models (SVM) and 2) models considering specific graph motifs do not seem to significantly improve performance. We also identify other graph kernels and GNN frameworks that show promise in motivating further research in brain network analysis.&lt;/p></description></item><item><title>A Tensor SVD-based Classification Algorithm Applied to fMRI Data</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/keegan-et-al-2021/</link><pubDate>Sun, 31 Oct 2021 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/keegan-et-al-2021/</guid><description>&lt;p>To analyze the abundance of multidimensional data, tensor-based frameworks have been developed. Traditionally, the matrix singular value decomposition (SVD) is used to extract the most dominant features from a matrix containing the vectorized data. While the SVD is highly useful for data that can be appropriately represented as a matrix, this step of vectorization causes us to lose the high-dimensional relationships intrinsic to the data. To facilitate efficient multidimensional feature extraction, we utilize a projection-based classification algorithm using the t-SVDM, a tensor analog of the matrix SVD. Our work extends the t-SVDM framework and the classification algorithm, both initially proposed for tensors of order 3, to any number of dimensions. We then apply this algorithm to a classification task using the StarPlus fMRI dataset. Our numerical experiments demonstrate that there exists a superior tensor-based approach to fMRI classification than the best possible equivalent matrix-based approach. Our results illustrate the advantages of our chosen tensor framework, provide insight into beneficial choices of parameters, and could be further developed for classification of more complex imaging data. We provide our Python implementation in this &lt;a href="https://github.com/elizabethnewman/tensor-fmri" target="_blank" rel="noopener">repository&lt;/a>.&lt;/p></description></item><item><title>Comparison of atlas-based and neural-network-based semantic segmentation for DENSE MRI images</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/buser-et-al-2021/</link><pubDate>Wed, 29 Sep 2021 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/buser-et-al-2021/</guid><description>&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/tdjXj3JdpQU" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen>&lt;/iframe>
&lt;p>Two segmentation methods, one atlas-based and one neural-network-based, were compared to see how well they can each automatically segment the brain stem and cerebellum in Displacement Encoding with Stimulated Echoes Magnetic Resonance Imaging (DENSE-MRI) data. The segmentation is a pre-requisite for estimating the average displacements in these regions, which have recently been proposed as biomarkers in the diagnosis of Chiari Malformation type I (CMI). In numerical experiments, the segmentations of both methods were similar to manual segmentations provided by trained experts. It was found that, overall, the neural-network-based method alone produced more accurate segmentations than the atlas-based method did alone, but that a combination of the two methods &amp;ndash; in which the atlas-based method is used for the segmentation of the brain stem and the neural-network is used for the segmentation of the cerebellum &amp;ndash; may be the most successful.&lt;/p></description></item><item><title>Alternating Minimization for Computed Tomography with Unknown Geometry Parameters</title><link>http://www.math.emory.edu/site/cmds-reuret/publications/pham-huynh-et-al-2021/</link><pubDate>Wed, 18 Aug 2021 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/publications/pham-huynh-et-al-2021/</guid><description>&lt;iframe width="560" height="315" src="https://www.youtube.com/embed/qdcGe9MKCoI" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture" allowfullscreen>&lt;/iframe>
&lt;p>Due to the COVID-19 pandemic, there is an increasing demand for portable CT machines worldwide in order to diagnose patients in a variety of settings [16]. This has lead to a need for CT image reconstruction algorithms that can produce high-quality images in the case when multiple types of geometry parameters have been perturbed. In this paper, we present an alternating descent algorithm to address this issue, where one step minimizes a regularized linear least squares problem, and the other minimizes a bounded non-linear least-square problem. Additionally, we survey existing methods to accelerate the convergence algorithm and discuss implementation details through the use of MATLAB packages such as IRtools and imfil. Finally, numerical experiments are conducted to show the effectiveness of our algorithm.&lt;/p></description></item><item><title/><link>http://www.math.emory.edu/site/cmds-reuret/apply/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/apply/</guid><description>&lt;h1 id="information-for-applicants">Information for Applicants&lt;/h1>
&lt;p>Our site provides several funded summer research opportunities for undergraduate students.
For summer 2026, applications can be submitted via &lt;a href="https://etap.nsf.gov/award/7517/opportunity/11727" target="_blank" rel="noopener">ETAP&lt;/a>.
We will begin reviewing applications on February 1 and all applications received by March 1 will receive full consideration. Until we finalize the roster (usually by early April) we will not decline qualified applicants.&lt;/p>
&lt;h2 id="eligibility">Eligibility&lt;/h2>
&lt;p>Students must be enrolled in an undergraduate program at a US institution during the summer and must not graduate before the end of the program.&lt;/p>
&lt;p>Participants must be US citizens or permanent residents due to NSF funding requirements.&lt;/p>
&lt;p>We expect participants to be fully available in-person during our REU phase and on campus for eight weeks during the summer. This year&amp;rsquo;s program will be held from June 15 until August 7.&lt;/p>
&lt;h2 id="expectations-and-pre-requisites">Expectations and Pre-Requisites&lt;/h2>
&lt;p>The research projects are designed to be accessible to participants with a strong working knowledge in Linear Algebra, Vector Calculus, Differential Equations and elementary programming experience. This can either be from past courses, courses students are enrolled in during the Spring term, or from independent studies. To help participants learn other project-specific materials, our site includes a comprehensive research training plan and our mentors are experienced and accessible to explain more advanced materials.&lt;/p>
&lt;p>Our activities include professional development, a weekly lunch seminar, and social dinners. Past participants have also often organized trips in the Atlanta area.&lt;/p>
&lt;h2 id="stipend-information">Stipend Information&lt;/h2>
&lt;p>Supported students will receive a stipend of 5,600 USD, travel support of up to 800 USD to and from Emory, and free on-campus accommodations.&lt;/p>
&lt;h2 id="housing">Housing&lt;/h2>
&lt;p>Participants will be housed on Emory’s beautiful Clairmont Campus. You can find a full description of the facilities &lt;a href="https://sihp.emory.edu/housing/index.html" target="_blank" rel="noopener">here&lt;/a>. We share an application link in April so everyone can register on time. Ideally, participants should arrive on the Sunday before the start of the program and move out on the Saturday after the program.&lt;/p>
&lt;p>Participants will generally be housed in 4-bedroom apartments. Every participant will have their own bedroom and bathrooms, kitchen, and living room will be shared. The SIHP will take care of room assignments. As much as possible, you will live with other students from this site but there may be one or two apartments that are mixed with summer students from other programs.&lt;/p></description></item><item><title/><link>http://www.math.emory.edu/site/cmds-reuret/contact/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/contact/</guid><description/></item><item><title/><link>http://www.math.emory.edu/site/cmds-reuret/overview/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/overview/</guid><description>&lt;h1 id="about-our-reu-site">About our REU site&lt;/h1>
&lt;p>The Emory Research Experience for Undergraduates site focuses on computational mathematics and its applications in data science, which impacts nearly every field of science, industry, and society.
Despite its growing importance, the number of academic training opportunities in data science has not kept pace with the rapid growth in demand from private and public entities.&lt;/p>
&lt;p>We have been running this program since the summer of 2021. Between 2021 and 2023, the site was held in conjunction with the Emory Research Experience for Teacher site and we hosted about a dozen in-service high school teachers.&lt;/p>
&lt;p>Our site emphasizes developing research and professional skills that enable participants to understand, conduct, and effectively communicate research in this booming area.
Each summer, our site trains at least twelve undergraduates for eight weeks.
The participants are mentored by faculty members from the &lt;a href="http://math.emory.edu/home/" target="_blank" rel="noopener">Department of Mathematics&lt;/a> and the &lt;a href="http://cs.emory.edu/home/" target="_blank" rel="noopener">Department of Computer Science&lt;/a>.&lt;/p>
&lt;p>We recruit students nationwide.&lt;/p>
&lt;h2 id="how-does-it-work">How does it work?&lt;/h2>
&lt;p>Our REU site introduces participants to the mathematical theory and computational tools used in applications.
Our projects include various topics ranging from data assimilation to machine learning.
Guided by faculty mentors, undergraduate students work toward creating new mathematical insights and designing practical solutions.&lt;/p>
&lt;p>Our site&amp;rsquo;s activities will be centered around a common theme that differs each year:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>2024&lt;/strong>: Learning from Images&lt;/li>
&lt;li>&lt;strong>2025&lt;/strong>: Combining Models and Data&lt;/li>
&lt;li>&lt;strong>2026&lt;/strong>: Discovering new Mathematics with Computational and Machine Learning Methods&lt;/li>
&lt;/ul>
&lt;p>Within each theme, faculty mentors will pose at least four research problems tackled by one student team.&lt;/p>
&lt;p>New insights of relevance to the broader scientific community will be created and disseminated in several ways, for example,&lt;/p>
&lt;ul>
&lt;li>oral presentations&lt;/li>
&lt;li>poster presentations&lt;/li>
&lt;li>student/teacher-authored publications&lt;/li>
&lt;li>open source software&lt;/li>
&lt;li>blog posts / project websites&lt;/li>
&lt;li>course materials for classroom-use&lt;/li>
&lt;/ul>
&lt;h2 id="what-pre-requisites-are-needed">What pre-requisites are needed?&lt;/h2>
&lt;p>By definition, the research projects will take participants beyond standard coursework.
To fill the gap, our site&amp;rsquo;s educational component introduces the participants to a range of mathematical techniques, including machine learning, deep neural networks, numerical linear algebra, optimization, partial differential equations, and statistics.
The faculty mentors also provide their mentees with professional and computational skills, including scientific writing, oral and poster presentations, and cloud computing.
The weekly seminar will feature group activities and faculty-led presentations on data and ethics, algorithmic bias, public scholarship.&lt;/p>
&lt;h2 id="information-for-applicants">Information for Applicants&lt;/h2>
&lt;p>Interested in joining? Learn more about about requirements, deadlines, and application materials on the &lt;a href="../apply/">information for applicants&lt;/a> page.&lt;/p></description></item><item><title/><link>http://www.math.emory.edu/site/cmds-reuret/readme/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/readme/</guid><description>&lt;h2 id="contents-emory-cmds-reuret-website">Contents: Emory CMDS REU/RET Website&lt;/h2>
&lt;p>This directory and its subdirectories contain the content that appears on our website.
Each subdirectory corresponds to a page on the website. Below we give a quick overview over the folders that you might want to edit.&lt;/p>
&lt;h2 id="home">home&lt;/h2>
&lt;p>This folder contains the elements of our main website. It mostly draws content from other places (such as featured projects) but also contains some text that can be edited in welcome.md&lt;/p>
&lt;h2 id="authors">authors&lt;/h2>
&lt;p>The subdirectories in this folder contain information about the people involved on our site. If you see a folder with your name, you can maintain your picture and bio here. If you are new or do not have a folder already,
you can copy one existing folder that corresponds to your role (e.g, mentor, student, teacher, &amp;hellip;)&lt;/p>
&lt;h2 id="projects">projects&lt;/h2>
&lt;p>This folder contains the website for the individual projects. The participants will have to edit their folder and can add images, videos, PDFs, etc.&lt;/p></description></item><item><title>People</title><link>http://www.math.emory.edu/site/cmds-reuret/people/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/people/</guid><description/></item><item><title>Summer 2021</title><link>http://www.math.emory.edu/site/cmds-reuret/summer2021/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/summer2021/</guid><description/></item><item><title>Summer 2022</title><link>http://www.math.emory.edu/site/cmds-reuret/summer2022/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/summer2022/</guid><description/></item><item><title>Summer 2023</title><link>http://www.math.emory.edu/site/cmds-reuret/summer2023/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/summer2023/</guid><description/></item><item><title>Summer 2024</title><link>http://www.math.emory.edu/site/cmds-reuret/summer2024/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/summer2024/</guid><description/></item><item><title>Summer 2025</title><link>http://www.math.emory.edu/site/cmds-reuret/summer2025/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/summer2025/</guid><description/></item><item><title>Summer 2026</title><link>http://www.math.emory.edu/site/cmds-reuret/summer2026/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>http://www.math.emory.edu/site/cmds-reuret/summer2026/</guid><description/></item></channel></rss>