Mittuniversitetet

miun.sePublikationer
Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Bottleneck-Aware Optimization of Distributed CNN Inference for Edge–Cloud IoT Systems
Mittuniversitetet, Fakulteten för naturvetenskap, teknik och medier, Institutionen för data- och elektroteknik (2023-).ORCID-id: 0000-0002-9903-1338
2026 (Engelska)Doktorsavhandling, sammanläggning (Övrigt vetenskapligt)
Abstract [en]

The proliferation of the Internet of Things (IoT) necessitates deploying Deep Learning (DL) models, specifically Convolutional Neural Networks (CNNs), on resource constrained edge devices. However, the high computational and memory demands of CNNs often exceed the capabilities of IoT nodes, while traditional cloud offloading suffers from latency and bandwidth limitations. This thesis proposes a comprehensive framework for Split Computing, enabling efficient distributed inference by partitioning CNNs between IoT nodes and edge servers. The core contribution is a bottleneck-aware feature compression mechanism designed to minimize data traffic at the partition point. The research demonstrates that combining partitioning with extreme quantization (down to 1-bit) and compression reduces data transmission by up to 99% with minimal accuracy loss. This approach is augmented by a novel hybrid structured pruning criterion, utilizing L2-norm magnitude and entropy, which selectively removes non informative channels to achieve significant speed-ups and energy savings compared to baseline execution modes. To address quantization induced accuracy degradation, the thesis introduces Time-Dependent Clustering Loss (TCL), a regularization technique that clusters activations during training to ensure robustness against extreme quantization errors. Furthermore,thecomplexselectionofpartitionpoints,compression ratios, and quantization levels is automated via CO-NAS (Compression Optimization Neural Architecture Search), a differentiable architecture search framework that efficiently discovers Pareto-optimal configurations. Validated on diverse hardware platforms (e.g., Raspberry Pi, NVIDIA Jetson) and standard datasets (CIFAR-100, TinyImageNet), these methodologies establish a robust pathway for Edge Intelligence. By unifying partitioning, quantization, compression, and automated search, this work provides a scalable solution for deploying high performance vision models in resource constrained IoT environments.

Ort, förlag, år, upplaga, sidor
Sundsvall: Mid Sweden University , 2026. , s. 59
Serie
Mid Sweden University doctoral thesis, ISSN 1652-893X ; 451
Nationell ämneskategori
Datorseende och lärande system
Identifikatorer
URN: urn:nbn:se:miun:diva-57326ISBN: 978-91-90017-64-7 (tryckt)OAI: oai:DiVA.org:miun-57326DiVA, id: diva2:2059121
Disputation
2026-06-09, M102, Holmgatan 10, Sundsvall, 09:00 (Engelska)
Opponent
Handledare
Anmärkning

Vid tidpunkten för disputationen var följande delarbete opublicerat: delarbete 5 inskickat.

At the time of the doctoral defence the following paper was unpublished: paper 5 submitted.

Tillgänglig från: 2026-05-12 Skapad: 2026-05-11 Senast uppdaterad: 2026-05-12Bibliografiskt granskad
Delarbeten
1. Waist Tightening of CNNs: A Case study on Tiny YOLOv3 for Distributed IoT Implementations
Öppna denna publikation i ny flik eller fönster >>Waist Tightening of CNNs: A Case study on Tiny YOLOv3 for Distributed IoT Implementations
Visa övriga...
2023 (Engelska)Ingår i: ACM International Conference Proceeding Series, Association for Computing Machinery (ACM), 2023, s. 241-246Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Computer vision systems in sensor nodes of the Internet of Things (IoT) based on Deep Learning (DL) are demanding because the DL models are memory and computation hungry while the nodes often come with tight constraints on energy, latency, and memory. Consequently, work has been done to reduce the model size or distribute part of the work to other nodes. However, then the question arises how these approaches impact the energy consumption at the node and the inference time of the system. In this work, we perform a case study to explore the impact of partitioning a Convolutional Neural Network (CNN) such that one part is implemented on the IoT node, while the rest is implemented on an edge device. The goal is to explore how the choice of partition point, quantization method and communication technology affects the IoT system. We identify possible partitioning points between layers, where we transform the feature maps passed between layers by applying quantization and compression to reduce the data sent over the communication channel between the two partitions in Tiny YOLOv3. The results show that a reduction of transmitted data by 99.8% reduces the network accuracy by 3 percentage points. Furthermore, the evaluation of various IoT communication protocols shows that the quantization of data facilitates CNN network partitioning with significant reduction of overall latency and node energy consumption. 

Ort, förlag, år, upplaga, sidor
Association for Computing Machinery (ACM), 2023
Nyckelord
CNN partitioning, convolutional neural networks, intelligence partitioning, Internet of Things, smart camera
Nationell ämneskategori
Kommunikationssystem
Identifikatorer
urn:nbn:se:miun:diva-48419 (URN)10.1145/3576914.3587518 (DOI)001054880600042 ()2-s2.0-85159789406 (Scopus ID)9798400700491 (ISBN)
Konferens
2023 Cyber-Physical Systems and Internet-of-Things Week, CPS-IoT Week 2023, 9 May 2023 through 12 May 2023
Tillgänglig från: 2023-06-07 Skapad: 2023-06-07 Senast uppdaterad: 2026-05-11Bibliografiskt granskad
2. Optimizing the IoT Performance: A Case Study on Pruning a Distributed CNN
Öppna denna publikation i ny flik eller fönster >>Optimizing the IoT Performance: A Case Study on Pruning a Distributed CNN
Visa övriga...
2023 (Engelska)Ingår i: 2023 IEEE Sensors Applications Symposium (SAS), 2023Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Implementing Convolutional Neural Networks (CNN) based computer vision algorithms in Internet of Things (IoT) sensor nodes can be difficult due to strict computational, memory, and latency constraints. To address these challenges, researchers have utilized techniques such as quantization, pruning, and model partitioning. Partitioning the CNN reduces the computational burden on an individual node, but the overall system computational load remains constant. Additionally, communication energy is also incurred. To understand the effect of partitioning and pruning on energy and latency, we conducted a case study using a feet detection application realized with Tiny Yolo-v3 on a 12th Gen Intel CPU with NVIDIA GeForce RTX 3090 GPU. After partitioning the CNN between the sequential layers, we apply quantization, pruning, and compression and study the effects on energy and latency. We analyze the extent to which computational tasks, data, and latency can be reduced while maintaining a high level of accuracy. After achieving this reduction, we offloaded the remaining partitioned model to the edge node. We found that over 90% computation reduction and over 99% data transmission reduction are possible while maintaining mean average precision above 95%. This results in up to 17x energy savings and up to 5.2x performance speed-up. 

Nyckelord
CNN, IoT, Partitioning, Pruning, Quantization, Tiny YOLO-v3
Nationell ämneskategori
Data- och informationsvetenskap
Identifikatorer
urn:nbn:se:miun:diva-49648 (URN)10.1109/SAS58821.2023.10254054 (DOI)001086399500037 ()2-s2.0-85174060733 (Scopus ID)9798350323078 (ISBN)
Konferens
2023 IEEE Sensors Applications Symposium, SAS 2023
Tillgänglig från: 2023-10-24 Skapad: 2023-10-24 Senast uppdaterad: 2026-05-11Bibliografiskt granskad
3. TCL: Time-dependent Clustering Loss for Optimizing Post-Training Feature Map Quantization for Partitioned DNNs
Öppna denna publikation i ny flik eller fönster >>TCL: Time-dependent Clustering Loss for Optimizing Post-Training Feature Map Quantization for Partitioned DNNs
Visa övriga...
2025 (Engelska)Ingår i: IEEE Access, E-ISSN 2169-3536, Vol. 13, s. 103640-103648Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

This paper introduces an enhanced approach for deploying deep learning models on resource-constrained IoT devices by combining model partitioning, autoencoder-based compression, quantization with Time Dependent Clustering Loss (TCL) regularization, and lossless compression, to reduce communication overhead, minimizing latency while maintaining accuracy. The autoencoder compresses feature maps at the partitioning point before quantization, effectively reducing data size and preserving accuracy. TCL regularization clusters activations at the partitioning point to align with quantization levels, minimizing quantization error and ensuring accuracy even with extreme low-bitwidth quantization. Our method is evaluated on classification models (ResNet-50, EfficientNetV2-S) and an object detection model (YOLOv10n) using the TinyImageNet-200 and Pascal VOC datasets. Deployed on Raspberry Pi 4 B and GPU, each model is tested across various partitioning points, quantization bit-widths (1-bit, 2-bit, and 3-bit), communication datarate (1MB/s to 10MB/s), and LZMA lossless compression. For a partitioned ResNet-50 after the convolutional stem block, the speed-up against a server solution is 2.33× and 1.85x compared to the all-in-node solution, with only a minimal accuracy drop of less than one percentage points. The proposed framework offers a scalable solution for deploying high-performance AI models on IoT devices, extending the feasibility of real-time inference in resource-constrained environments. 

Ort, förlag, år, upplaga, sidor
Institute of Electrical and Electronics Engineers (IEEE), 2025
Nyckelord
CNN, IoT, Partitioning, Quantization
Nationell ämneskategori
Datorsystem
Identifikatorer
urn:nbn:se:miun:diva-54756 (URN)10.1109/ACCESS.2025.3579107 (DOI)001512606800010 ()2-s2.0-105008273568 (Scopus ID)
Tillgänglig från: 2025-06-24 Skapad: 2025-06-24 Senast uppdaterad: 2026-05-11
4. Efficient Edge Inference via Entropy and Magnitude-Aware Feature Map Pruning in Partitioned CNNs
Öppna denna publikation i ny flik eller fönster >>Efficient Edge Inference via Entropy and Magnitude-Aware Feature Map Pruning in Partitioned CNNs
Visa övriga...
2025 (Engelska)Ingår i: 2025 International Conference on Machine Learning and Applications (ICMLA), IEEE conference proceedings, 2025, s. 1028-1033Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Edge–to–cloud inference for vision based models often stalls on the activation data that must cross the link between an IoT device and its server. This paper integrates Time-dependent Clustering Loss (TCL) with a lightweight, layer-specific novel hybrid L2 + Entropy channel-pruning rule applied exactly at the partition boundary. TCL aligns feature maps to discrete levels, enabling aggressive post-training quantization, while the hybrid pruning criterion removes up to 90 % of channels without noticeable accuracy loss.Experiments on ResNet50, EfficientNetV2-S, and YOLOv10n running on a Jetson Orin Nano node demonstrates clear gains across three realistic uplink capacities. At 100 Mbit s−1, partitioned inference completes in 1.2 ms (ResNet50), 1.1 ms (EfficientNetV2-S) and 2.8 ms (YOLOv10n), yielding approximately 6×, 7× and 18× while speed-ups over an all-server baseline still surpassing full on-device execution. With uplinks of 300 Mbit s−1 and 500 Mbit s−1, the communication penalty shrinks, yet the proposed pipeline retains more than 2× acceleration over all-edge processing and limits accuracy degradation to a 1–2 % point drop.Because pruning is confined to a single layer and relies only on post-training activation statistics, integration is simple and incurs negligible overhead. The combined TCL quantization and L2 + Entropy pruning strategy therefore offers a practical, bandwidth-aware solution for deploying modern CNNs in IoT-edge scenarios.

Ort, förlag, år, upplaga, sidor
IEEE conference proceedings, 2025
Nationell ämneskategori
Data- och informationsvetenskap
Identifikatorer
urn:nbn:se:miun:diva-57324 (URN)10.1109/icmla66185.2025.00156 (DOI)2-s2.0-105036979343 (Scopus ID)
Konferens
2025 International Conference on Machine Learning and Applications (ICMLA), Boca Raton, FL, USA, 03-05 December, 2025
Forskningsfinansiär
KK-stiftelsenInterreg
Tillgänglig från: 2026-05-11 Skapad: 2026-05-11 Senast uppdaterad: 2026-05-13Bibliografiskt granskad

Open Access i DiVA

fulltext(14505 kB)194 nedladdningar
Filinformation
Filnamn FULLTEXT01.pdfFilstorlek 14505 kBChecksumma SHA-512
7c576690c788e27ed6e1ef8afad4918ee5a581c21592b58e621477fda87297002065518eebbc7f4186a72c476072f767cb561fdca9f8d92dc8c9ee2eca6de017
Typ fulltextMimetyp application/pdf

Person

Saqib, Eiraj

Sök vidare i DiVA

Av författaren/redaktören
Saqib, Eiraj
Av organisationen
Institutionen för data- och elektroteknik (2023-)
Datorseende och lärande system

Sök vidare utanför DiVA

GoogleGoogle Scholar
Antalet nedladdningar är summan av nedladdningar för alla fulltexter. Det kan inkludera t.ex tidigare versioner som nu inte längre är tillgängliga.

isbn
urn-nbn

Altmetricpoäng

isbn
urn-nbn
Totalt: 220 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf