Skip to main navigation Skip to search Skip to main content

Cardio-AI-ReAccel: reconfigurable accelerators for artificial intelligence in cardiology

  • Muhammad Shakeel Akram

Student thesis: Doctoral Thesis

Abstract

This thesis advances the design of resource-efficient, privacy-preserving, and high performance machine learning frameworks for real-time cardiac diagnosis on edge devices. With the growing integration of smart sensors, embedded AI, and connected wearables, eHealth is transforming cardiovascular care. However, challenges such as limited hardware resources, latency constraints, privacy concerns, and the need for on-device continual learning hinder practical deployment, especially as arrhythmias remain diagnostically complex and a leading cause of mortality worldwide.

To address these challenges, the research proposes a suite of hardware-software co-designed solutions combining quantized deep neural networks (qDNNs), federated learning, and field-programmable gate array (FPGA)-based acceleration. First, an ultra-lightweight embedded DNN using pruning was deployed on Arduino Nano BLE 33 Sense, achieving94% accuracy in MIT-BIH arrhythmia classification with a model size below 876 KB compared to the state of the art, demonstrating the feasibility of edge-based real-time screening. Building on this, the thesis introduces mixed-precision qDNN (MPQ-DNN)accelerators for continual learning on FPGAs, enabling on-device training without costlyre-synthesis. These accelerators balance diagnostic accuracy, power, and memory through adaptive bitwidth tuning and streaming-friendly designs. They achieved 93.71% top-1accuracy, 13.82 KB model size, 1545 throughput, and 18.4 µs latency, surpassing existing works by up to 144×, while reducing the accelerator’s weight update latency by a factor of(Continual Learning Cycles–1)×.

To support privacy-compliant learning, the fpgaDPFL framework was developed as the first end-to-end design space exploration tool for differentially private federated learning(DPFL) on FPGAs. While no prior workflow supports full DPFL training on reconfigurable hardware, fpgaDPFL supports end-to-end DPFL training, including back propagation, quantized gradient perturbation, knowledge distillation, and secure aggregation, enabling low-latency, private, decentralized learning. Its real-world evaluation on federated ECG classification confirms its potential for secure collaboration across data-isolated institutions, allowing trade-off exploration between accuracy, energy, differential privacy budgets, and hardware utilization.

Further contributions include qMLP, achieving 80.1% accuracy with memory footprint of 186 KB for low-correlation cardiac classes, and FINN-Custom-Verification, a design-time verification framework exposing transformation issues in FINN when using binary activations with zero-padding. A 1-bit Activation_Resolver was proposed, yielding a 12.7 KB binarized CNN (BCNN) with 81.65% accuracy. This demonstrates the feasibility of deploying accurate cardiac classifiers on ultra-constrained edge devices such as wearables, where memory, power, and storage are critically limited. This opens new avenues for real-time, always on cardiac monitoring within the TinyML paradigm, supporting continual learning and DPFL under strict resource budgets. Additionally, an evolutionary hardware-in-the-loop (EvoHIL) framework was introduced, perturbing FPGA weights post-deployment to activate underutilized neurons, boosting BCNN accuracy to 85% without retraining or re-synthesis.To optimize accelerator throughput, the Aopt algorithm was developed to align folding factors with memory interface widths, reducing padding overhead by 96.77%, improving throughput by 33.2%, and cutting runtime by 25% compared to FINN.

Finally, this work reveals a novel vulnerability in DPFL-secure systems: internal power-based free-rider attacks. Using the proposed EvoWeight strategy, it demonstrates how adversaries can craft shared gradients that stealthily increase power consumption (+17.41%),latency (+2.31%) and reduce throughput (-4.75%) without degrading the accuracy of the model. This can also pose serious risks such as thermal stress, denial-of-service (DoS), and hardware degradation or failure in critical edge deployments.

Together, these contributions form a holistic edge-AI design methodology that integrates adaptive quantization, continual and federated learning, and secure FPGA acceleration, empowering real-time, energy-efficient, and privacy-aware cardiac diagnostics for next-generation eHealth systems.

This thesis is embargoed until 31 October 2027

Date of AwardOct 2025
Original languageEnglish
SupervisorSharatchandra Varma Bogaraju (Supervisor)

Keywords

  • DNN
  • TinyMML
  • EdgeAI
  • privacy
  • federated learning
  • continual learning
  • FPGA accelerators
  • cardiac diagnosis

Cite this

'