Classification and Pose Recovery from Image Sequences

01.01.2011 - 31.12.2013
Research funding project

Neuroscience studies have pointed out that a human's brain work with 3D representation of objects and somehow (unclear how) stores the 3D information to achieve object detection and recognition at the level we humans have been experiencing. In computer vision applications, 3D information is often gathered by using multiple cameras. When multiple cameras having an overlapping view and a wide baseline are used it is not possible to reconstruct an observed object based on corresponding points because such features can’t be found between an image showing the left side of an object and an image showing the right side of the object. Additionally, when using video streams as input, the trajectory of the observed object may not be consistent, because each frame is treated independently. The goal of this project is to develop a real time framework which overcomes these problems and presents a new way to classify both rigid and non-rigid objects and determine their 3D pose by using predefined 3D models in combination with videos as input data. In a first step, a rough classification and pose estimation is done by detecting some (e.g. 100) of the most likely poses of the observed object using synthetic 3D models for each frame of a video. In a second step the pose is refined over multiple frames (e.g. 30) using some predefined metrics (e.g. the pose must not change much and the class must not change at all between two temporally following frames). Additionally, the pose is refined over multiple distributed cameras, which should lead to robust and smooth results. Applications are located in the area of 3D surveillance networks and can mainly be used for - Urban traffic analysis: As the framework should be able to do a precise classification of different cars in terms of size, shape and colour, the algorithms can be used on parking lots, in garages etc. for counting, classifying and identifying specific vehicles and pedestrians. - Security scenarios: Abnormal behaviour of humans can be detected and alarms can be given to the security staff. Due to the massive amount of cameras, it is unlikely for humans to track multiple objects on multiple cameras simultaneously. In this area, automatic approaches would improve their work enormously.

People

Project leader

Project personnel

Institute

Grant funds

  • FFG - Österr. Forschungsförderungs- gesellschaft mbH (National) Austrian Research Promotion Agency (FFG)

Research focus

  • Media Informatics and Visual Computing: 100%

Keywords

GermanEnglish
3D Objekt Klassifikation3D object classification
3D Bildverstehen3D scene understanding
3D Vision3D Vision
Pose estimationPose estimation
bildfolgenimage sequences

External partner

  • CogVis Software und Consulting GmbH

Publications