This Blog is maintained by the Robot Perception and Learning lab at CSIE, NTU, Taiwan. Our scientific interests are driven by the desire to build intelligent robots and computers, which are capable of servicing people more efficiently than equivalent manned systems in a wide variety of dynamic and unstructured environments.
Monday, June 27, 2011
Lab meeting June 29th (Jim): A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Stephane Ross, Geoffrey Gordon, and J. Andrew (Drew) Bagnell
Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (AISTATS), April, 2011.
Abstracts:
Sequential prediction problems such as imitation learning, where future observations depend on previous predictions (actions), violate the common i.i.d. assumptions made in statistical learning. ... In this paper, we propose a new iterative algorithm, which trains a stationary deterministic policy, that can be seen as a no regret algorithm in an online learning setting. We show that any such no regret algorithm, combined with additional reduction assumptions, must find a policy with good performance under the distribution of observations it induces in such sequential settings.
Link
Monday, June 20, 2011
Lab Meeting June 22th (Chih-Chung):Minimum Snap Trajectory Generation and Control for Quadrotors (ICRA2011,best paper)
Authors: Daniel Mellinger and Vijay Kumar
Abstracts:
We address the controller design and the trajectory
generation for a quadrotor maneuvering in three
dimensions in a tightly constrained setting typical of indoor
environments. In such settings, it is necessary to allow for
significant excursions of the attitude from the hover state and
small angle approximations cannot be justified for the roll
and pitch. We develop an algorithm that enables the real-time
generation of optimal trajectories through a sequence of 3-D
positions and yaw angles, while ensuring safe passage through
specified corridors and satisfying constraints on velocities,
accelerations and inputs. A nonlinear controller ensures the
faithful tracking of these trajectories. Experimental results
illustrate the application of the method to fast motion (5-10
body lengths/second) in three-dimensional slalom courses.
[link]
Wednesday, June 15, 2011
Lab Meeting June 15th (Shao-Chen): Distributed Robust Data Fusion Based on Dynamic Voting (ICRA2011)
Authors: Eduardo Montijano, Sonia Mart´ınez and Carlos Sagues
Abstract:
Data association mistakes, estimation and measurement errors are some of the factors that can contribute to incorrect observations in robotic sensor networks. In order to act reliably, a robotic network must be able to fuse and correct its perception of the world by discarding any outlier information. This is a difficult task if the network is to be deployed remotely and the robots do not have access to groundtruth sites or manual calibration. In this paper, we present a novel, distributed scheme for robust data fusion in autonomous robotic networks. The proposed method adapts the RANSAC algorithm to exploit measurement redundancy, and enables robots determine an inlier observation with local communications. Different hypotheses are generated and voted for using a dynamic consensus algorithm. As the hypotheses are computed, the robots can change their opinion making the voting process dynamic. Assuming that at least one hypothesis is initialized with only inliers, we show that the method converges to the maximum likelihood of all the inlier observations in a general instance. Several simulations exhibit the good performance of the algorithm, which also gives acceptable results in situations where the conditions to guarantee convergence do not hold.
[link]
Tuesday, June 14, 2011
Lab Meeting June 15th (David): Sparse Scene Flow Segmentation for Moving Object Detection (Intelligent Vehicles Symposium 2011)
Authors: P. Lenz, J. Ziegler, A. Geiger, M. Roser
Abstract:
Modern driver assistance systems such as collision avoidance or intersection assistance need reliable information on the current environment. Extracting such information from camera-based systems is a complex and challenging task for inner city taffic scenarios. This paper presents an approach for object detection utilizing sparse scene flow. For consecutive stereo images taken from a moving vehicle, corresponding interest points are extracted. Thus, for every interest point, disparity and optical flow values are known and consequently, scene flow can be calculated. Adjacent interest points describing a similar scene flow are considered to belong to one rigid object. The proposed method does not rely on object classes and allows for a robust detection of dynamic objects in traffic scenes. Leading vehicles are continuously detected for several frames. Oncoming objects are detected within five frames after their appearance.
Link: http://www.rainsoft.de/publications/iv11b.pdf
Tuesday, June 07, 2011
Lab Meeting June 8th, 2011 (Jeff): Incremental Construction of the Saturated-GVG for Multi-Hypothesis Topological SLAM
Topological SLAM
Authors: Tong Tao, Stephen Tully, George Kantor, and Howie Choset
Abstract:
The generalized Voronoi graph (GVG) is a topological representation of an environment that can be incrementally constructed with a mobile robot using sensor-based control. However, because of sensor range limitations, the GVG control law will fail when the robot moves into a large open area. This paper discusses an extended GVG approach to topological navigation and mapping: the saturated generalized Voronoi graph (S-GVG), for which the robot employs an additional wall-following behavior to navigate along obstacles at the range limit of the sensor. In this paper, we build upon previous work related to the S-GVG and provide two important contributions: 1) a rigorous discussion of the control laws and algorithm modifications that are necessary for incremental construction of the S-GVG with a mobile robot, and 2) a method for incorporating the S-GVG into a novel multi-hypothesis SLAM algorithm for loop-closing and localization. Experiments with a wheeled mobile robot in an office-like environment validate the ffectiveness of the proposed approach.
Link:
IEEE International Conference on Robotics and Automation(ICRA), 2011
http://www.cs.cmu.edu/~biorobotics/papers/icra11_tao.pdf
LocalLink
Wednesday, June 01, 2011
Lab Meeting June 1, 2011 (Alan): Semantic Structure from Motion (CVPR 2011)
Tuesday, May 31, 2011
Lab Meeting June 1, 2011 (Wang Li): Articulated pose estimation with flexible mixtures-of-parts (CVPR 2011)
Yi Yang
Deva Ramanan
Abstract
We describe a method for human pose estimation in static images based on a novel representation of part models. Notably, we do not use articulated limb parts, but rather capture orientation with a mixture of templates for each part. We describe a general, flexible mixture model for capturing contextual co-occurrence relations between parts, augmenting standard spring models that encode spatial relations. We show that such relations can capture notions of local rigidity. When co-occurrence and spatial relations are tree-structured, our model can be efficiently optimized with dynamic programming. We present experimental results on standard benchmarks for pose estimation that indicate our approach is the state-of-the-art system for pose estimation, outperforming past work by 50% while being orders of magnitude faster.
Paper Link
Monday, May 16, 2011
ICRA 2011 Awards
- WINNER! Characterization of Oscillating Nano Knife for Single Cell Cutting by Nanorobotic Manipulation System Inside ESEM: Yajing Shen, Masahiro Nakajima, Seiji Kojima, Michio Homma, Yasuhito Ode, Toshio Fukuda [pdf]
- Wireless Manipulation of Single Cells Using Magnetic Microtransporters: Mahmut Selman Sakar, Edward Steager, Anthony Cowley, Vijay Kumar, George J Pappas
- Hierarchical Planning in the Now: Leslie Kaelbling, Tomas Lozano-Perez
- Selective Injection and Laser Manipulation of Nanotool Inside a Specific Cell Using Optical Ph Regulation and Optical Tweezers: Hisataka Maruyama, Naoya Inoue, Taisuke Masuda, Fumihito Arai
- Configuration-Based Optimization for Six Degree-Of-Freedom Haptic Rendering for Fine Manipulation: Dangxiao Wang, Xin Zhang, Yuru Zhang, Jing Xiao
Best Vision Paper
- Model-Based Localization of Intraocular Microrobots for Wireless Electromagnetic Control: Christos Bergeles, Bradley Kratochvil, Bradley J. Nelson
- Fusing Optical Flow and Stereo in a Spherical Depth Panorama Using a Single-Camera Folded Catadioptric Rig: Igor Labutov, Carlos Jaramillo, Jizhong Xiao
- 3-D Scene Analysis Via Sequenced Predictions Over Points and Regions: Xuehan Xiong, Daniel Munoz, James Bagnell, Martial Hebert
- Fast and Accurate Computation of Surface Normals from Range Images: Hernan Badino, Daniel Huber, Yongwoon Park, Takeo Kanade
- WINNER! Sparse Distance Learning for Object Recognition Combining RGB and Depth Information: Kevin Lai, Liefeng Bo, Xiaofeng Ren, Dieter Fox [pdf]
Best Automation Paper
- WINNER! Automated Cell Manipulation: Robotic ICSI: Zhe Lu, Xuping Zhang, Clement Leung, Navid Esfandiari, Robert Casper, Yu Sun [pdf]
- Efficient AUV Navigation Fusing Acoustic Ranging and Side-Scan Sonar: Maurice Fallon, Michael Kaess, Hordur Johannsson, John Leonard
- Vision-Based 3D Bicycle Tracking Using Deformable Part Model and Interacting Multiple Model Filter: Hyunggi Cho, Paul E. Rybski, Wende Zhang
- High-Accuracy GPS and GLONASS Positioning by Multipath Mitigation Using Omnidirectional Infrared Camera: Taro Suzuki, Mitsunori Kitamura, Yoshiharu Amano, Takumi Hashizume
- Deployment of a Point and Line Feature Localization System for an Outdoor Agriculture Vehicle: Jacqueline Libby, George Kantor
Best Medical Robotics Paper
- Design of Adjustable Constant-Force Forceps for Robot-Assisted Surgical Manipulation: Chao-Chieh Lan, Jung-Yuan Wang
- Design Optimization of Concentric Tube Robots Based on Task and Anatomical Constraints: Chris Bedell, Jesse Lock, Andrew Gosline, Pierre Dupont
- GyroLock - First in Vivo Experiments of Active Heart Stabilization Using Control Moment Gyro (CMG): Julien Gagne, Olivier Piccin, Edouard Laroche, Michele Diana, Jacques Gangloff
- Metal MEMS Tools for Beating-Heart Tissue Approximation: Evan Butler, Chris Folk, Adam Cohen, Nikolay Vasilyev, Rich Chen, Pedro del Nido, Pierre Dupont
- WINNER! An Articulated Universal Joint Based Flexible Access Robot for Minimally Invasive Surgery: Jianzhong Shang, David Noonan, Christopher Payne, James Clark, Mikael Hans Sodergren, Ara Darzi, Guang-Zhong Yang [pdf]
Best Conference Paper
- WINNER! Minimum Snap Trajectory Generation and Control for Quadrotors: Daniel Mellinger, Vijay Kumar [pdf]
- Autonomous Multi-Floor Indoor Navigation with a Computationally Constrained Micro Aerial Vehicle: Shaojie Shen, Nathan Michael, Vijay Kumar
- Dexhand : A Space Qualfied Multi-Fingered Robotic Hand: Maxime Chalon, Armin Wedler, Andreas Baumann, Wieland Bertleff, Alexander Beyer, Jörg Butterfass, Markus Grebenstein, Robin Gruber, Franz Hacker, Erich Krämer, Klaus Landzettel, Maximilian Maier, Hans-Juergen Sedlmayr, Nikolaus Seitz, Fabian Wappler, Bertram Willberg, Thomas Wimboeck, Frederic Didot, Gerd Hirzinger
- Time Scales and Stability in Networked Multi-Robot Systems: Mac Schwager, Nathan Michael, Vijay Kumar, Daniela Rus
- Bootstrapping Bilinear Models of Robotic Sensorimotor Cascades: Andrea Censi, Richard Murray
KUKA Service Robotics Best Paper
- Distributed Coordination and Data Fusion for Underwater Search: Geoffrey Hollinger, Srinivas Yerramalli, Sanjiv Singh, Urbashi Mitra, Gaurav Sukhatme
- WINNER! Dynamic Shared Control for Human-Wheelchair Cooperation: Qinan Li, Weidong Chen, Jingchuan Wang [pdf]
- Towards Joint Attention for a Domestic Service Robot -- Person Awareness and Gesture Recognition Using Time-Of-Flight Cameras: David Droeschel, Jorg Stuckler, Dirk Holz, Sven Behnke
- Electromyographic Evaluation of Therapeutic Massage Effect Using Multi-Finger Robot Hand: Ren C. Luo, Chih-Chia Chang
Best Video
- Catching Flying Balls and Preparing Coffee: Humanoid Rollin'Justin Performs Dynamic and Sensitive Tasks: Berthold Baeuml, Florian Schmidt, Thomas Wimboeck, Oliver Birbach, Alexander Dietrich, Matthias Fuchs, Werner Friedl, Udo Frese, Christoph Borst, Markus Grebenstein, Oliver Eiberger, Gerd Hirzinger
- Recent Advances in Quadrotor Capabilities: Daniel Mellinger, Nathan Michael, Michael Shomin, Vijay Kumar
- WINNER! High Performance of Magnetically Driven Microtools with Ultrasonic Vibration for Biomedical Innovations: Masaya Hagiwara, Tomohiro Kawahara, Lin Feng, Yoko Yamanishi, Fumihito Arai [pdf]
Best Cognitive Robotics Paper
- WINNER! Donut As I Do: Learning from Failed Demonstrations: Daniel Grollman, Aude Billard [pdf]
- A Discrete Computational Model of Sensorimotor Contingencies for Object Perception and Control of Behavior: Alexander Maye, Andreas Karl Engel
- Skill Learning and Task Outcome Prediction for Manipulation: Peter Pastor, Mrinal Kalakrishnan, Sachin Chitta, Evangelos Theodorou, Stefan Schaal
- Integrating Visual Exploration and Visual Search in Robotic Visual Attention: The Role of Human-Robot Interaction: Momotaz Begum, Fakhri Karray
Tuesday, May 03, 2011
Lab Meeting May 3rd (Andi): Face/Off: Live Facial Puppetry
Monday, May 02, 2011
Lab Meeting May 3( KuenHan ), Multiple Targets Tracking in World Coordinate with a Single, Minimally Calibrated Camera (ECCV 2010)
Author: Wongun Choi, Silvio Savarese.
Abstract:
Tracking multiple objects is important in many application
domains. We propose a novel algorithm for multi-object tracking that
is capable of working under very challenging conditions such as min-
imal hardware equipment, uncalibrated monocular camera, occlusions
and severe background clutter. To address this problem we propose a
new method that jointly estimates object tracks, estimates correspond-
ing 2D/3D temporal trajectories in the camera reference system as well
as estimates the model parameters (pose, focal length, etc) within a
coherent probabilistic formulation. Since our goal is to estimate stable
and robust tracks that can be univocally associated to the object IDs,
we propose to include in our formulation an interaction (attraction and
repulsion) model that is able to model multiple 2D/3D trajectories in
space-time and handle situations where objects occlude each other. We
use a MCMC particle ltering algorithm for parameter inference and
propose a solution that enables accurate and e cient tracking and cam-
era model estimation. Qualitative and quantitative experimental results
obtained using our own dataset and the publicly available ETH dataset
shows very promising tracking and camera estimation results.
Link
Website
Wednesday, April 20, 2011
NTU PAL Thesis Defense: Mobile Robot Localization in Large-scale Dynamic Environments
- Thesis draft: http://any.csie.ntu.edu.tw/thesis/yang_thesis-v1_0.pdf
Sunday, April 17, 2011
Lab Meeting April 20, 2011 (fish60): Donut as I do: Learning from failed demonstrations
Tuesday, April 12, 2011
Lab Meeting April 13, 2011 (Will): Hilbert Space Embeddings of Hidden Markov Models (ICML2010)
Lab Meeting April 13, 2011 (Jimmy): WiFi-SLAM Using Gaussian Process Latent Variable Models (IJCAI2007)
In: IJCAI 2007
Authors: Brian Ferris, Dieter Fox, and Neil Lawrence
Abstract
WiFi localization, the task of determining the physical location of a mobile device from wireless signal strengths, has been shown to be an accurate method of indoor and outdoor localization and a powerful building block for location-aware applications. However, most localization techniques require a training set of signal strength readings labeled against a ground truth location map, which is prohibitive to collect and maintain as maps grow large. In this paper we propose a novel technique for solving the WiFi SLAM problem using the Gaussian Process Latent Variable Model (GPLVM) to determine the latent-space locations of unlabeled signal strength data. We show how GPLVM, in combination with an appropriate motion dynamics model, can be used to reconstruct a topological connectivity graph from a signal strength sequence which, in combination with the learned Gaussian Process signal strength model, can be used to perform efficient localization.
[pdf]
Tuesday, March 29, 2011
Lab Meeting March 30, 2011 (Chih-Chung): Progress Report
Lab Meeting March 30, 2011 (Chung-Han): Progress Report
Tuesday, March 22, 2011
Lab Meeting March 23, 2011 (David): Object detection and tracking for autonomous navigation in dynamic environments (IJRR 2010)
Authors: Andreas Ess, Konrad Schindler, Bastian Leibe, Luc Van Gool
Abstract:
We address the problem of vision-based navigation in busy inner-city locations, using a stereo rig mounted on a mobile platform. In this scenario semantic information becomes important: rather than modeling moving objects as arbitrary obstacles, they should be categorized and tracked in order to predict their future behavior. To this end, we combine classical geometric world mapping with object category detection and tracking. Object-category-specific detectors serve to find instances of the most important object classes (in our case pedestrians and cars). Based on these detections, multi-object tracking recovers the objects' trajectories, thereby making it possible to predict their future locations, and to employ dynamic path planning. The approach is evaluated on challenging, realistic video sequences recorded at busy inner-city locations.
Link
Lab Meeting March 23, 2011 (Shao-Chen): A Comparison of Track-to-Track Fusion Algorithms for Automotive Sensor Fusion (MFI2008)
Authors: Stephan Matzka and Richard Altendorfer
Abstract:
In exteroceptive automotive sensor fusion, sensor data are usually only available as processed, tracked object data and not as raw sensor data. Applying a Kalman filter to such data leads to additional delays and generally underestimates the fused objects' covariance due to temporal correlations of individual sensor data as well as inter-sensor correlations. We compare the performance of a standard asynchronous Kalman filter applied to tracked sensor data to several algorithms for the track-to-track fusion of sensor objects of unknown correlation, namely covariance union, covariance intersection, and use of cross-covariance. For the simulation setup used in this paper, covariance intersection and use of cross-covariance turn out to yield significantly lower errors than a Kalman filter at a comparable computational load.
Link
Monday, March 14, 2011
Lab Meeting March 16th, 2011 (Andi): 3D Deformable Face Tracking with a Commodity Depth Camera
Wednesday, March 09, 2011
Lab Meeting March 9th, 2011(KuoHuei): progress report
Tuesday, March 08, 2011
Lab Meeting March 9, 2011 (Wang Li): Real-time Identification and Localization of Body Parts from Depth Images (ICRA 2010)
Christian Plagemann
Varun Ganapathi
Daphne Koller
Sebastian Thrun
Abstract
We deal with the problem of detecting and identifying body parts in depth images at video frame rates. Our solution involves a novel interest point detector for mesh and range data that is particularly well suited for analyzing human shape. The interest points, which are based on identifying geodesic extrema on the surface mesh, coincide with salient points of the body, which can be classified using local shape descriptors. Our approach also provides a natural way of estimating a 3D orientation vector for a given interest point. This can be used to normalize the local shape descriptors to simplify the classification problem as well as to directly estimate the orientation of body parts in space.
Experiments show that our interest points in conjunction with a boosted patch classifier are significantly better in detecting body parts in depth images than state-of-the-art sliding-window based detectors.
Paper Link
Thursday, March 03, 2011
Article: Perception beyond the Here and Now
Computer, February 2011, pp. 86–88
A multitude of senses provide us with information about the here and now. What we see, hear, and feel in turn shape how we perceive our surroundings and understand the world. Our senses are extremely limited, however, and ever since humans began creating and using technology, they have tried to enhance their natural perception in various ways. (pdf)
Monday, February 28, 2011
Lab Meeting March 2nd, 2011 (Jeff): Observability-based Rules for Designing Consistent EKF SLAM Estimators
Authors: Guoquan P. Huang, Anastasios Mourikis, and Stergios I. Roumeliotis
Abstract:
In this work, we study the inconsistency problem of extended Kalman filter (EKF)-based simultaneous localization and mapping (SLAM) from the perspective of observability. We analytically prove that when the Jacobians of the process and measurement models are evaluated at the latest state estimates during every time step, the linearized error-state system employed in the EKF has an observable subspace of dimension higher than that of the actual, non-linear, SLAM system. As a result, the covariance estimates of the EKF undergo reduction in
directions of the state space where no information is available, which is a primary cause of the inconsistency. Based on these theoretical results, we propose a general framework for improving the consistency of EKF-based SLAM. In this framework, the EKF linearization points are selected in a way that ensures that the resulting linearized system model has an observable subspace of appropriate dimension. We describe two algorithms that are instances of this paradigm. In the first, termed observability constrained (OC)-EKF, the linearization points are selected so as to minimize their expected errors (i.e. the difference between the linearization point and the true state) under the observability constraints. In the second, the filter Jacobians are calculated using the first-ever available estimates for all state variables. This latter approach is termed first-estimates Jacobian (FEJ)-EKF. The proposed algorithms have been tested both in simulation and experimentally, and are shown to significantly outperform the standard EKF both in terms of accuracy and consistency.
Link:
The International Journal of Robotics Research(IJRR), Vol.5 April 2010
http://ijr.sagepub.com/content/29/5/502.full.pdf+html
Sunday, February 20, 2011
Wednesday, February 09, 2011
Lab Meeting February 14, 2011 (fish60): Feature Construction for Inverse Reinforcement Learning
Sergey Levine, Zoran Popović, Vladlen Koltun
NIPS 2010
Abstract:
The goal of inverse reinforcement learning is to find a reward function for a
Markov decision process, given example traces from its optimal policy. Current
IRL techniques generally rely on user-supplied features that form a concise basis
for the reward. We present an algorithm that instead constructs reward features
from a large collection of component features, by building logical conjunctions of
those component features that are relevant to the example policy. Given example
traces, the algorithm returns a reward function as well as the constructed features.
Link
Lab Meeting February 14, 2011 (Alan): Multibody Structure-from-Motion in Practice (PAMI 2010)
Authors: Kemal Egemen Ozden, Konrad Schindler, and Luc Van Gool
Abstract—Multibody structure from motion (SfM) is the extension of classical SfM to dynamic scenes with multiple rigidly moving objects. Recent research has unveiled some of the mathematical foundations of the problem, but a practical algorithm which can handle realistic sequences is still missing. In this paper, we discuss the requirements for such an algorithm, highlight theoretical issues and practical problems, and describe how a static structure-from-motion framework needs to be extended to handle real dynamic scenes. Theoretical issues include different situations in which the number of independently moving scene objects changes: Moving objects can enter or leave the field of view, merge into the static background (e.g., when a car is parked), or split off from the background and start moving independently. Practical issues arise due to small freely moving foreground objects with few and short feature tracks. We argue that all of these difficulties need to be handled online as structure-from-motion estimation progresses, and present an exemplary solution using the framework of probabilistic model-scoring.
Link
Monday, January 17, 2011
Lab Meeting January 17( KuenHan ), Moving Object Detection by Multi-View Geometric Techniques from a Single Camera Mounted Robot (IROS 2009)
Author: Abhijit Kundu, K Madhava Krishna and Jayanthi Sivaswamy
Abstract:
The ability to detect, and track multiple moving
objects like person and other robots, is an important prerequisite
for mobile robots working in dynamic indoor environments.
We approach this problem by detecting independently moving
objects in image sequence from a monocular camera mounted
on a robot. We use multi-view geometric constraints to classify
a pixel as moving or static. The first constraint, we use, is the
epipolar constraint which requires images of static points to
lie on the corresponding epipolar lines in subsequent images.
In the second constraint, we use the knowledge of the robot
motion to estimate a bound in the position of image pixel along
the epipolar line. This is capable of detecting moving objects
followed by a moving camera in the same direction, a so-called
degenerate configuration where the epipolar constraint fails.
To classify the moving pixels robustly, a Bayesian framework
is used to assign a probability that the pixel is stationary
or dynamic based on the above geometric properties and
the probabilities are updated when the pixels are tracked in
subsequent images. The same framework also accounts for the
error in estimation of camera motion. Successful and repeatable
detection and pursuit of people and other moving objects in
realtime with a monocular camera mounted on the Pioneer
3DX, in a cluttered environment confirms the efficacy of the
method.
Link
Sunday, January 09, 2011
Lab Meeting January 10th, 2011(Jimmy) : Accurate Image Localization Based on Google Maps Street View (ECCV 2010)
Authors: Amir Roshan Zamir, Mubarak Shah
In ECCV 2010
Abstract
Finding an image's exact GPS location is a challenging computer vision problem that has many real-world applications. In this paper, we address the problem of fi nding the GPS location of images with an accuracy which is comparable to hand-held GPS devices. We leverage a structured data set of about 100,000 images build from Google Maps Street View as the reference images. We propose a localization method in which the SIFT descriptors of the detected SIFT interest points in the reference images are indexed using a tree. In order to localize a query image, the tree is queried using the detected SIFT descriptors in the query image. A novel GPS-tag-based pruning method removes the less reliable descriptors. Then, a smoothing step with an associated voting scheme is utilized; this allows each query descriptor to vote for the location its nearest neighbor belongs to, in order to accurately localize the query image. A parameter called Confidence of Localization which is based on the Kurtosis of the distribution of votes is de fined to determine how reliable the localization of a particular image is. In addition, we propose a novel approach to localize groups of images accurately in a hierarchical manner. First, each image is localized individually; then, the rest of the images in the group are matched against images in the neighboring area of the found first match. The fi nal location is determined based on the Confidence of Localization parameter. The proposed image group localization method can deal with very unclear queries which are not capable of being geolocated individually.
[pdf]
Monday, January 03, 2011
Lab Meeting January 3rd, 2011(Will) : Neural Prothesis & Realtime Bayes Tracking
Neural prothesis is a field that use brain to control motors to help disable people.
I'll report my survey on the neural prothesis decoding algorithm.
Po-Wei
Sunday, December 26, 2010
Lab Meeting January 3rd, 2011(David) :Vision-Based Behavior Prediction in Urban Traffic Environments by Scene Categorization (BMVC 2010)
Authors: Martin Heracles, Fernando Martinelli and Jannik Fritsch
Abstract:
We propose a method for vision-based scene understanding in urban traffic environments that predicts the appropriate behavior of a human driver in a given visual scene. The method relies on a decomposition of the visual scene into its constituent objects by image segmentation and uses segmentation-based features that represent both their identity and spatial properties. We show how the behavior prediction can be naturally formulated as scene categorization problem and how ground truth behavior data for learning a classifier can be automatically generated from any monocular video sequence recorded from a moving vehicle, using structure from motion techniques. We evaluate our method both quantitatively and qualitatively on the recently proposed CamVid dataset, predicting the appropriate velocity and yaw rate of the car as well as their appropriate change for both day and dusk sequences. In particular, we investigate the impact of the underlying segmentation and the number of behavior classes on the quality of these predictions
link
Wednesday, December 22, 2010
Lab Meeting December 27, 2010(Chih Chung) : Lozano-Perez. Belief space planning assuming maximum likelihood observations.(RSS 2010)
Authors:Robert Platt Jr., Russ Tedrake, Leslie Kaelbling, Tomas Lozano-Perez
Abstract:
We cast the partially observable control problem as
a fully observable underactuated stochastic control problem in
belief space and apply standard planning and control techniques.
One of the difficulties of belief space planning is modeling the
stochastic dynamics resulting from unknown future observations.
The core of our proposal is to define deterministic beliefsystem
dynamics based on an assumption that the maximum
likelihood observation (calculated just prior to the observation)
is always obtained. The stochastic effects of future observations
are modelled as Gaussian noise. Given this model of the dynamics,
two planning and control methods are applied. In the first, linear
quadratic regulation (LQR) is applied to generate policies in the
belief space. This approach is shown to be optimal for linear-
Gaussian systems. In the second, a planner is used to find locally
optimal plans in the belief space. We propose a replanning
approach that is shown to converge to the belief space goal
in a finite number of replanning steps. These approaches are
characterized in the context of a simple nonlinear manipulation
problem where a planar robot simultaneously locates and grasps
an object.
link
Sunday, December 19, 2010
Lab Meeting December 20, 2010(Chung-Han): progress report
Sunday, December 12, 2010
Lab Meeting December 13, 2010(ShaoChen): DDF-SAM: Fully Distributed SLAM using Constrained Factor Graphs(IROS2010)
Authors: Alexander Cunningham, Manohar Paluri, and Frank Dellaert
Abstract:
We address the problem of multi-robot distributed SLAM with an extended Smoothing and Mapping (SAM) approach to implement Decentralized Data Fusion (DDF). We present DDF-SAM, a novel method for efficiently and robustly distributing map information across a team of robots, to achieve scalability in computational cost and in communication bandwidth and robustness to node failure and to changes in network topology. DDF-SAM consists of three modules: (1) a local optimization module to execute single-robot SAM and condense the local graph; (2) a communication module to collect and propagate condensed local graphs to other robots, and (3) a neighborhood graph optimizer module to combine local graphs into maps describing the neighborhood of a robot. We demonstrate scalability and robustness through a simulated example, in which inference is consistently faster than a comparable naive approach.
[link]
Monday, December 06, 2010
Lab Meeting December 6th, 2010(Nicole): Acoustic Source Localization and Tracking Using Track Before Detect
(IEEE TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING, 2010)
Authors: Maurice F. Fallon, Simon Godsill
Abstract:
Particle Filter-based Acoustic Source Localization algorithms attempt to track the position of a sound source—one or more people speaking in a room—based on the current data from a microphone array as well as all previous data up to that point. This paper first discusses some of the inherent behavioral traits of the steered beamformer localization function. Using conclusions drawn from that study, a multitarget methodology for acoustic source tracking based on the Track Before Detect (TBD) framework is introduced. The algorithm also implicitly evaluates source activity using a variable appended to the state vector. Using the TBD methodology avoids the need to identify a set of source measurements and also allows for a vast increase in the number of particles used for a comparitive computational load which results in increased tracking stability in challenging recording environments. An evaluation of tracking performance is given using a set of real speech recordings with two simultaneously active speech sources.
[link]
Lab Meeting December 6th, 2010(KuoHuei): progress report
Sunday, November 28, 2010
Lab Meeting November 29, 2010 (Wang Li): Adaptive Pose Priors for Pictorial Structures (CVPR 2010)
Benjamin Sapp
Chris Jordan
Ben Taskar
Abstract
The structure and parameterization of a pictorial structure model is often restricted by assuming tree dependency structure and unimodal, data-independent pairwise interactions, which fail to capture important patterns in the data. On the other hand, local methods such as kernel density estimation provide nonparametric flexibility but require large amounts of data to generalize well. We propose a simple semi-parametric approach that combines the tractability of pictorial structure inference with the flexibility of non-parametric methods by expressing a subset of model parameters as kernel regression estimates from a learned sparse set of exemplars. This yields query-specific, image-dependent pose priors. We develop an effective shape-based kernel for upper-body pose similarity and propose a leave-one-out loss function for learning a sparse subset of exemplars for kernel regression. We apply our techniques to two challenging datasets of human figure parsing and advance the state-of-the-art (from 80% to 86% on the Buffy dataset), while using only 15% of the training data as exemplars.
Paper Link
Saturday, November 27, 2010
Lab Meeting November 29th, 2010 (Jeff): Sub-Meter Indoor Localization in Unmodified Environments with Inexpensive Sensors
Authors: Morgan Quigley, David Stavens, Adam Coates, and Sebastian Thrun
Abstract:
The interpretation of uncertain sensor streams for localization is usually considered in the context of a robot. Increasingly, however, portable consumer electronic devices, such as smartphones, are equipped with sensors including WiFi radios, cameras, and inertial measurement units (IMUs). Many tasks typically associated with robots, such as localization, would be valuable to perform on such devices. In this paper, we present an approach for indoor localization exclusively using the low-cost sensors typically found on smartphones. Environment modification is not needed. We rigorously evaluate our method using ground truth acquired using a laser range scanner. Our evaluation includes overall accuracy and a comparison of the contribution of individual sensors. We find experimentally that fusion of multiple sensor modalities is necessary for optimal performance and demonstrate sub-meter localization accuracy.
Link:
IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS), October 2010
http://www-cs.stanford.edu/people/dstavens/iros10/quigley_etal_iros10.pdf
or
local_copy
Video:
http://www.cs.stanford.edu/people/dstavens/iros10/quigley_etal_iros10.mp4
Monday, November 22, 2010
Lab Meeting November 22, 2010 (Andi): Three-Dimensional Mapping with Time-of-Flight Cameras
Sunday, November 21, 2010
Lab Meeting November 22, 2010 (Alan): Temporary Maps for Robust Localization in Semi-static Environments (IROS 2010)
Monday, November 15, 2010
Lab Meeting November 15( KuenHan ), 3D Reconstruction of a Moving Point from a Series of 2D Projections (ECCV 2010)
Author: Hyun Soo Park, Takaaki Shiratori, Iain Matthews, and Yaser Sheikh
Abstract
This paper presents a linear solution for reconstructing the 3D trajectory of a moving point from its correspondence in a collection of 2D perspective images, given the 3D spatial pose and time of capture of the cameras that produced each image. Triangulation-based solutions do not apply, as multiple views of the point may not exist at each instant in time. A geometric analysis of the problem is presented and a criterion, called reconstructibility, is defined to precisely characterize the cases when reconstruction is possible, and how accurate it can be. We apply the linear reconstruction algorithm to reconstruct the time evolving 3D structure of several real-world scenes, given a collection of non-coincidental 2D images.
LinkSunday, November 14, 2010
Lab Meeting November 15, 2010 (fish60): Unfreezing the Robot: Navigation in Dense, Interacting Crowds
Author: Peter Trautman and Andreas Krause
Abstract—In this paper, we study the safe navigation of a mobile robot through crowds of dynamic agents with uncertain trajectories. Existing algorithms suffer from the “freezing robot” problem: once the environment surpasses a certain level of complexity, the planner decides that all forward paths are unsafe, and the robot freezes in place (or performs unnecessary aneuvers) to avoid collisions. ... In this work, we demonstrate that both the individual prediction and the predictive uncertainty have little to do with the frozen robot problem. Our key insight is that dynamic agents solve the frozen robot problem by engaging in “joint collision avoidance”: They cooperatively make room to create feasible trajectories. We develop IGP, a nonparametric statistical model based on Dependent Output Gaussian Processes that can estimate crowd interaction from data. Our model naturally captures the non-Markov nature of agent trajectories, as well as their goal-driven navigation. We then show how planning in this model can be efficiently implemented using particle based inference.
Link
Monday, November 01, 2010
CMU PhD Thesis Defense: Geolocation with Range: Robustness, Efficiency and Scalability
Joseph A. Djugash
Geolocation with Range: Robustness, Efficiency and Scalability
November 05, 2010, 10:00 a.m., NSH 1507
Abstract
This thesis explores the topic of geolocation with range. A robust method for localization and SLAM (Simultaneous Localization and Mapping) is proposed. This method uses a polar parameterization of the state to achieve accurate estimates of the nonlinear and multi-modal distributions in range-only systems. Several experimental evaluations on real robots reveal the reliability of this method.
Scaling such a system to large network of nodes, increases the computational load on the system due to the increased state vector. To alleviate this problem, we propose the use of a distributed estimation algorithm based on the belief propagation framework. This method distributes the estimation task, such that each node only estimates its local network, greatly reducing the computation performed by any individual node. However, the method does not provide any guarantees on the convergence of its solution in general graphs. Convergence is only guaranteed for non-cyclic graphs (ie. trees). Thus, an extension of this approach which reduces any arbitrary graph to a spanning tree is presented. This enables the proposed decentralized localization method to provide guarantees on its convergence.
[LINK][PDF]
Thesis Committee
Sanjiv Singh, Chair
George Kantor
Howie Choset
Wolfram Burgard, University of Freiburg
Sunday, October 31, 2010
Lab Meeting November 1, 2010 (Will): Visual Event Recognition in Videos by Learning from Web Data (CVPR 2010)
Author: Lixin Duan, Dong Xu, Ivor W. Tsang, Jiebo Luo
Abstract:
We propose a visual event recognition framework for consumer domain videos by leveraging a large amount of loosely labeled web videos (e.g., from YouTube). First, we propose a new aligned space-time pyramid matching method to measure the distances between two video clips, where each video clip is divided into space-time volumes over multiple levels. We calculate the pairwise distances between any two volumes and further integrate the information from different volumes with Integer-flow Earth Mover’s Distance (EMD) to explicitly align the volumes. Second, we propose a new cross-domain learning method in order to 1) fuse the information from multiple pyramid levels and features (i.e., space-time feature and static SIFT feature) and 2) cope with the considerable variation in feature dis- tributions between videos from two domains (i.e., web do- main and consumer domain). For each pyramid level and each type of local features, we train a set of SVM classifiers based on the combined training set from two domains using multiple base kernels of different kernel types and parameters, which are fused with equal weights to obtain an average classifier. Finally, we propose a cross-domain learning method, referred to as Adaptive Multiple Kernel Learning (A-MKL), to learn an adapted classifier based on multiple base kernels and the prelearned average classifiers by minimizing both the structural risk functional and the mismatch between data distributions from two domains. Extensive experiments demonstrate the effectiveness of our proposed framework that requires only a small number of labeled consumer videos by leveraging web data.
Friday, October 29, 2010
Lab meeting Nov. 01 2010, (Chih-Chung) POMDPs for robotic tasks with mixed observability (RSS 2009)
Author:Sylvie C.W.Ong, Shao Wei Png, David Hsu and Wee Sun Lee.
Abstract:
Partially observable Markov decision processes
(POMDPs) provide a principled mathematical framework for
motion planning of autonomous robots in uncertain and dynamic
environments. They have been successfully applied to
various robotic tasks, but a major challenge is to scale up
POMDP algorithms for more complex robotic systems. Robotic
systems often have mixed observability: even when a robot’s
state is not fully observable, some components of the state
may still be fully observable. Exploiting this, we use a factored
model to represent separately the fully and partially observable
components of a robot’s state and derive a compact lowerdimensional
representation of its belief space. We then use this
factored representation in conjunction with a point-based algorithm
to compute approximate POMDP solutions. Separating
fully and partially observable state components using a factored
model opens up several opportunities to improve the efficiency
of point-based POMDP algorithms. Experiments show that on
standard test problems, our new algorithm is many times faster
than a leading point-based POMDP algorithm.
Thursday, October 28, 2010
News: University of Chicago, Cornell Researchers Develop Universal Robotic Gripper

Robotic hands are usually just that -- hands -- but some researchers from the University of Chicago and Cornell University (with a little help from iRobot) have taken a decidedly different approach for their so-called universal robotic gripper. As you can see above, the gripper is actually a balloon that can conform to and grip just about any small object, and hang onto it firmly enough to pick it up. What's the secret? After much testing, the researchers found that ground coffee was the best substance to fill the balloon with -- to grab an object, the gripper simply creates a vacuum in the balloon (much like a vacuum-sealed bag of coffee), and it's then able to let go of the object just by releasing the vacuum. Simple, but it works. Head on past the break to check it out in action. [via engadget]
Monday, October 25, 2010
Lab meeting Oct. 25 2010, (David) Threat-aware Path Planning in Uncertain Urban Environments (IROS 2010)
Authors: Georges S. Aoude, Brandon D. Luders, Daniel S. Levine, and Jonathan P. How
Abstract:
This paper considers the path planning problem
for an autonomous vehicle in an urban environment populated
with static obstacles and moving vehicles with uncertain intents.
We propose a novel threat assessment module, consisting of
an intention predictor and a threat assessor, which augments
the host vehicle’s path planner with a real-time threat value
representing the risks posed by the estimated intentions of
other vehicles. This new threat-aware planning approach is
applied to the CL-RRT path planning framework, used by the
MIT team in the 2007 DARPA Grand Challenge. The strengths
of this approach are demonstrated through simulation and
experiments performed in the RAVEN testbed facilities
[local copy]
[link ]
[local video]
[video]
Monday, October 11, 2010
Lab meeting Oct. 11 2010, (Shao-Chen) Consistent data association in multi-robot systems with limited communications(RSS 2010)
Authors: Rosario Aragues,Eduardo Montijano, and Carlos Sagues
Abstract:
In this paper we address the data association
problem of features observed by a robot team with limited communications.
At every time instant, each robot can only exchange
data with a subset of the robots, its neighbors. Initially, each
robot solves a local data association with each of its neighbors.
After that, the robots execute the proposed algorithm to agree
on a data association between all their local observations which
is globally consistent. One inconsistency appears when chains of
local associations give rise to two features from one robot being
associated among them. The contribution of this work is the
decentralized detection and resolution of these inconsistencies.
We provide a fully decentralized solution to the problem. This
solution does not rely on any particular communication topology.
Every robot plays the same role, making the system robust to
individual failures. Information is exchanged exclusively between
neighbors. In a finite number of iterations, the algorithm finishes
with a data association which is free of inconsistent associations.
In the experiments, we show the performance of the algorithm
under two scenarios. In the first one, we apply the resolution
and detection algorithm for a set of stochastic visual maps. In
the second, we solve the feature matching between a set of images
taken by a robotic team.
[link]
Lab meeting Oct. 11th 2010, (Nicole) Improvement in Listening Capability for Humanoid Robot HRP-2(ICRA 2010)
Title: Improvement in Listening Capability for Humanoid Robot HRP-2 (ICRA2010)
Authors: Toru Takahashi, Kazuhiro Nakadai, Kazunori Komatani, Tetsuya Ogata and Hiroshi G. Okuno.
Abstract:
This paper describes improvement of sound source separation for a simultaneous automatic speech recognition (ASR) system of a humanoid robot. A recognition error in the system is caused by a separation error and interferences of other sources. In separability, an original geometric source separation (GSS) is improved. Our GSS uses a measured robot’s head related transfer function (HRTF) to estimate a separation matrix. As an original GSS uses a simulated HRTF calculated based on a distance between microphone and sound source, there is a large mismatch between the simulated and the measured transfer functions. The mismatch causes a severe degradation of recognition performance.
Faster convergence speed of separation matrix reduces separation error. Our approach gives a nearer initial separation matrix based on a measured transfer function from an optimal separation matrix than a simulated one. As a result, we expect that our GSS improves the convergence speed. Our GSS is also able to handle an adaptive step-size parameter.
These new features are added into open source robot audition software (OSS) called”HARK” which is newly updated as version 1.0.0. The HARK has been installed on a HRP-2 humanoid with an 8-element microphone array. The listening capability of HRP-2 is evaluated by recognizing a target speech signal which is separated from a simultaneous speech signal by three talkers. The word correct rate (WCR) of ASR improves by 5 points under normal acoustic environments and by 10 points under noisy environments. Experimental results show that HARK 1.0.0 improves the robustness against noises.
Lab meeting Oct. 11th 2010, (Andi) Dynamic 3D Scene Analysis for Acquiring Articulated Scene Models
Sunday, October 10, 2010
News: Google's Self-Driving Cars
By Sebastian Thrun, Google
Read the full article.
Google Cars Drive Themselves, in Traffic
By John Markoff, The New York Times
Read the full article.
Friday, October 08, 2010
News: MIT Media Lab Medical Mirror
Sunday, October 03, 2010
Lab Meeting October 4th, 2010 (Jeff): Progress Report
Lab Meeting October 4th, 2010(KuoHuei): progress report
Monday, September 27, 2010
Lab Meeting September 27, 2010 (Wang Li): Monocular 3D Pose Estimation and Tracking by Detection (CVPR 2010)
Mykhaylo Andriluka
Stefan Roth
Bernt Schiele
Abstract
Automatic recovery of 3D human pose from monocular
image sequences is a challenging and important research
topic with numerous applications. Although current methods
are able to recover 3D pose for a single person in controlled
environments, they are severely challenged by realworld
scenarios, such as crowded street scenes. To address
this problem, we propose a three-stage process building on
a number of recent advances. The first stage obtains an initial
estimate of the 2D articulation and viewpoint of the person
from single frames. The second stage allows early data
association across frames based on tracking-by-detection. The third and
final stage uses those tracklet-based estimates as robust image
observations to reliably recover 3D pose. We demonstrate
state-of-the-art performance on the HumanEva II
benchmark, and also show the applicability of our approach
to articulated 3D tracking in realistic street conditions.
Paper Link