Mittuniversitetet

miun.sePublikationer
Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Processing chain for 3D histogram of gradients based real-time object recognition
Mittuniversitetet, Fakulteten för naturvetenskap, teknik och medier, Institutionen för elektronikkonstruktion.
Mittuniversitetet, Fakulteten för naturvetenskap, teknik och medier, Institutionen för elektronikkonstruktion.
Mittuniversitetet, Fakulteten för naturvetenskap, teknik och medier, Institutionen för elektronikkonstruktion.
2021 (Engelska)Ingår i: International Journal of Advanced Robotic Systems, ISSN 1729-8806, E-ISSN 1729-8814, Vol. 18, nr 1, artikel-id 1729881420978363Artikel i tidskrift (Refereegranskat) Published
Abstract [en]

3D object recognition has been a cutting-edge research topic since the popularization of depth cameras. These cameras enhance the perception of the environment and so are particularly suitable for autonomous robot navigation applications. Advanced deep learning approaches for 3D object recognition are based on complex algorithms and demand powerful hardware resources. However, autonomous robots and powered wheelchairs have limited resources, which affects the implementation of these algorithms for real-time performance. We propose to use instead a 3D voxel-based extension of the 2D histogram of oriented gradients (3DVHOG) as a handcrafted object descriptor for 3D object recognition in combination with a pose normalization method for rotational invariance and a supervised object classifier. The experimental goal is to reduce the overall complexity and the system hardware requirements, and thus enable a feasible real-time hardware implementation. This article compares the 3DVHOG object recognition rates with those of other 3D recognition approaches, using the ModelNet10 object data set as a reference. We analyze the recognition accuracy for 3DVHOG using a variety of voxel grid selections, different numbers of neurons (N-h ) in the single hidden layer feedforward neural network, and feature dimensionality reduction using principal component analysis. The experimental results show that the 3DVHOG descriptor achieves a recognition accuracy of 84.91% with a total processing time of 21.4 ms. Despite the lower recognition accuracy, this is close to the current state-of-the-art approaches for deep learning while enabling real-time performance.

Ort, förlag, år, upplaga, sidor
2021. Vol. 18, nr 1, artikel-id 1729881420978363
Nationell ämneskategori
Datorgrafik och datorseende
Identifikatorer
URN: urn:nbn:se:miun:diva-41626DOI: 10.1177/1729881420978363ISI: 000619537100001Scopus ID: 2-s2.0-85099946553OAI: oai:DiVA.org:miun-41626DiVA, id: diva2:1537169
Tillgänglig från: 2021-03-15 Skapad: 2021-03-15 Senast uppdaterad: 2025-09-25
Ingår i avhandling
1. Semi-Autonomous Navigation of Powered Wheelchairs: 2D/3D Sensing and Positioning Methods
Öppna denna publikation i ny flik eller fönster >>Semi-Autonomous Navigation of Powered Wheelchairs: 2D/3D Sensing and Positioning Methods
2021 (Engelska)Doktorsavhandling, sammanläggning (Övrigt vetenskapligt)
Abstract [en]

Autonomous driving and assistance systems have become a reality for the automotive industry to improve driving safety in the car. Hence, the cars use a variety of sensors, cameras and image processing techniques to measure their surroundings and control their direction, braking and speed for obstacle avoidance or autonomously driving applications.Like the automotive industry, powered wheelchairs also require safety systems to ensure their operation, especially when the user has controlling limitations, but also to develop new applications to improve its usability. One of the applications is focused on developing a new contactless control of a powered wheelchair using the position of a caregiver beside it as a control reference. Contactless control can prevent control errors, but it can also provide better and more equal communication between the wheelchair user and the caregiver

This thesis evaluates the camera requirements for a contactless powered wheelchair control and the 2D/3D image processing techniques for caregiver recognition and position measurement beside the powered wheelchair. The research evaluates the strength and limitations of different depth camera technologies for caregiver feet detection above the ground plane to select the proper camera for the application. Then, a hand-crafted 3D object descriptor is evaluated for caregiver feet recognition and compared with respect to a state-of-the-art deep learning object detector. Results for both methods are good, however, the hand-crafted descriptor suffers from segmentation errors and consequently, their accuracy is lower. After the depth camera and image processing techniques evaluation, results show that it is possible to use only an RGB camera to recognize and measure his or her relative position.

Ort, förlag, år, upplaga, sidor
Sundsvall: Mid Sweden University, 2021. s. 64
Serie
Mid Sweden University doctoral thesis, ISSN 1652-893X ; 354
Nyckelord
3D object recognition, YOLO, YOLO-Tiny, 3DHOG, Histogram-of-Oriented-Gradients, ModelNet40, Feature descriptor, Intel RealSense, Depth camera, Wheelchair
Nationell ämneskategori
Datorgrafik och datorseende
Identifikatorer
urn:nbn:se:miun:diva-43829 (URN)978-91-89341-32-6 (ISBN)
Disputation
2021-12-09, O102, Mittuniversitetet, Sundsvall, 16:06 (Engelska)
Opponent
Handledare
Anmärkning

Vid tidpunkten för disputationen var följande delarbeten opublicerade: delarbete 5 inskickat.

At the time of the doctoral defence the following papers were unpublished: paper 5 submitted.

Tillgänglig från: 2021-11-24 Skapad: 2021-11-23 Senast uppdaterad: 2025-09-25Bibliografiskt granskad

Open Access i DiVA

fulltext(1024 kB)929 nedladdningar
Filinformation
Filnamn FULLTEXT01.pdfFilstorlek 1024 kBChecksumma SHA-512
4e3d4bb3dea43373102933dc83c9051b7cb96f0ec766a036b6db4d50181f7c1062e53d824f3c0d50b733dcef0634d765ceb136d246d29833a254074f0c7160b7
Typ fulltextMimetyp application/pdf

Övriga länkar

Förlagets fulltextScopus

Person

Vilar, CristianKrug, SilviaThörnberg, Benny

Sök vidare i DiVA

Av författaren/redaktören
Vilar, CristianKrug, SilviaThörnberg, Benny
Av organisationen
Institutionen för elektronikkonstruktion
I samma tidskrift
International Journal of Advanced Robotic Systems
Datorgrafik och datorseende

Sök vidare utanför DiVA

GoogleGoogle Scholar
Totalt: 930 nedladdningar
Antalet nedladdningar är summan av nedladdningar för alla fulltexter. Det kan inkludera t.ex tidigare versioner som nu inte längre är tillgängliga.

doi
urn-nbn

Altmetricpoäng

doi
urn-nbn
Totalt: 316 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf