This Blog is maintained by the Robot Perception and Learning lab at CSIE, NTU, Taiwan. Our scientific interests are driven by the desire to build intelligent robots and computers, which are capable of servicing people more efficiently than equivalent manned systems in a wide variety of dynamic and unstructured environments.
Thursday, August 20, 2009
Lab Meeting August 24, 2009(Chung-Han) : Monitoring an intersection using a network of laser scanners
Huijing Zhao; Jinshi Cui; Hongbin Zha; Katabira, K.; Xiaowei Shao; Shibasaki, R.
Intelligent Transportation Systems, 2008. ITSC 2008. 11th International IEEE Conference on12-15 Oct. 2008 Page(s):428 - 433
Abstract : In this research, a novel system for monitoring an intersection using a network of single-row laser range scanners (subsequently abbreviated as "laser scanner") is proposed. Laser scanners are set on the road side to profile an intersection horizontally from different viewpoints. This is done so that cross sections of the intersection are captured at a high scanning rate (e.g., 37 Hz) and to contain the contour points of the moving objects entering the intersection. Different laser scanners data are integrated into a common spatial-temporal coordinate system and processed. Thus, the moving objects inside the intersection are detected and tracked to estimate their state parameters, such as: location, speed, and direction at each time instance. An experiment was conducted in central Beijing, where six laser scanners were used to cover a three-way intersection. A digital copy of the dynamic intersection was measured, and, through data processing, a large quantity of physical dimension traffic data was obtained.
[link]
Wednesday, August 19, 2009
Lab Meeting August 24, 2009(Jimmy): One-Shot Learning of Object Categories
Thursday, August 13, 2009
Lab Meeting August 17, 2009 (Any): RANSAC-based DARCES
Title:Moving Obstacle Detection in Highly Dynamic Scenes(ICRA09)
Authors:A. Ess, B. Leibe, K. Schindler, and L. van Gool.
Abstract
We address the problem of vision-based multipersontracking in busy pedestrian zones using a stereo rigmounted on a mobile platform. Specifically, we are interestedin the application of such a system for supporting pathplanning algorithms in the avoidance of dynamic obstacles.
The complexity of the problem calls for an integrated solution, whichextracts as much visual information as possible and combinesit through cognitive feedback.
We propose such an approach,which jointly estimates camera position, stereo depth, objectdetections, and trajectories based only on visual information.The interplay between these components is represented in agraphical model. For each frame, we first estimate the groundsurface together with a set of object detections. Based onthese results, we then address object interactions and estimatetrajectories. Finally, we employ the tracking results to predictfuture motion for dynamic objects and fuse this informationwith a static occupancy map estimated from dense stereo.
The approach is experimentally evaluated on several longand challenging video sequences from busy inner-city locationsrecorded with different mobile setups. The results show thatthe proposed integration makes stable tracking and motionprediction possible, and thereby enables path planning incomplex and highly dynamic scenes.
link
webpage
NewScientist: Why humans can't navigate out of a paper bag
Magazine issue 2721.
[Full Article]
This is an interesting article mentioning that human beings do not perform metric-SLAM and may perform topological SLAM poorly. Take a look.
-Bob
Monday, August 10, 2009
IJCAI 2009 Talk: From Low-level Sensors to High-level Intelligence: Activity Recognition Links the Knowledge Food Chain
Author: Qiang Yang, The Hong Kong University of Science and Technology
Description
Sensors provide computer systems with a window to the outside world. Activity recognition "sees" what is in the window to predict the locations, trajectories, actions, goals and plans of humans and objects. Building an activity recognition system requires a full range of interaction from statistical inference on lower level sensor data to symbolic AI at higher levels, where prediction results and acquired knowledge are passed up each level to form a knowledge food chain. In this talk, I will give an overview of activity recognition and explore its relation to other fields, including planning and knowledge acquisition, machine learning and Web search. I will also describe its applications in assistive technologies, security monitoring and mobile commerce.
Link
Sunday, August 09, 2009
Lab Meeting August 10, 2009(Nicole):Learning Sound Location from a Single Microphone (ICRA 2009)
Title:Learning Sound Location from a Single Microphone (ICRA 2009)
Authors:AshutoshSaxena and AndrewY. Ng
Abstract:
We consider the problem of estimating the incident angle of a sound, using only a single microphone. The ability to perform monaural (single-ear) localization isimportant to many animals; indeed, monaural cues are also the primary method by which humans decide if a sound comes from the front or back, as well as estimate its elevation. Such monaural localization is made possible by the structure of the pinna (outer ear), which modifies sound in a way that is dependent on its incident angle. In this paper, we propose a machine learning approach to monaural localization, using only a single microphone and an “artificial pinna” (that distorts sound in a direction-dependent way). Our approach models the typical distribution of natural and artificial sounds, as well as the direction-dependent changes to sounds induced by the pinna. Our experimental results also show that the algorithm is able to fairly accurately localize a wide range of sounds, such as human speech, dog barking, waterfall, thunder, and so on. In contrast to microphone arrays, this approach also offers the potential of significantly more compact, as well as lower cost and power, devices for sounds localization.
[link]
Friday, August 07, 2009
Paper: Modeling Groups of Plausible Virtual Pedestrians
IEEE Computer Graphics and Applications, July/August 2009, pp. 54–63
Crowd simulation is enjoying considerable success in a number of applied domains, most notably in evacuation scenarios in which simulated crowd behaviors can help improve the safety of interior building designs. However, not all applications involving virtual populace have the overarching goal of realistic simulation. In many cases, it's necessary only that viewers perceive the crowd as realistic. In many of the latest movies or video games involving large numbers of virtual actors, liberties can be taken in displaying those far away or otherwise obscured from the eye, if this doesn't noticeably diminish the viewing experience. For example, such simulations can reduce the level of detail or forgo collision avoidance calculations to allow simulation of a larger crowd or enhanced behaviors for individuals deemed most likely to occupy viewers' attention. (Full PDF)
FRC Seminar Special Time:Information Sharing in Large Heterogeneous Teams, August 13, 2009
PhD Student
Robotics Institute
Carnegie Mellon University
Wednesday, August 05, 2009
MIT CSAIL Thesis Defense: Visual Sense Disambiguation: A Multimodal Approach
If a picture is worth a thousand words, can a thousand words be worth a training image? Most successful object recognition algorithms require manually annotated images of objects to be collected for training. The amount of human effort required to collect training data has limited most approaches to the several hundred object categories available in the labeled datasets. While human-annotated image data is scarce, additional sources of information can be used as weak labels, reducing the need for human supervision. In this thesis, we use three types of information to learn models of object categories: speech, text and dictionaries. We demonstrate that our use of non-traditional information sources facilitates automatic acquisition of visual object models for arbitrary words without requiring any labeled image examples.
Spoken object references occur in many scenarios: interaction with an assistant robot, voice-tagging of photos, etc. Existing reference resolution methods are unimodal, relying either only on image features, or only on speech recognition. We propose a method that uses both the image of the object and the speech segment referring to it to disambiguate the underlying object label. We show that even noisy speech input helps visual recognition, and vice versa. We also explore two sources of linguistic sense information: the words surrounding images on web pages, and dictionary entries for nouns that refer to objects. Keywords that index images on the web have been used as weak object labels, but these tend to produce noisy datasets with many unrelated images. We use unlabeled text, dictionary definitions, and semantic relations between concepts to learn a refined model of image sense. Our model can work with as little supervision as a single English word. We apply this model to a dataset of web images indexed by polysemous keywords, and show that it improves both retrieval of specific senses, and the resulting object classifiers.
Tuesday, August 04, 2009
Lab Meeting August 10, 2009(Gary):In Between 3D Active Appearance Models and 3D Morphable Models
Authors: Jingu Heo, Marios Savvides
Abstract:
In this paper we propose a novel method of generating
3D Morphable Models (3DMMs) from 2D images. We
develop algorithms of 3D face reconstruction from a
sparse set of points acquired from 2D images. In order
to establish correspondence between images precisely, we
Combined Active Shape Models (ASMs) and Active Appearance
Models (AAMs)(CASAAMs) in an intelligent way,
showing improved performance on pixel-level accuracy
and generalization to unseen faces. The CASAAMs are
applied to the images of different views of the same person
to extract facial shapes across pose. These 2D shapes are
combined for reconstructing a sparse 3D model. The point
density of the model is increased by the Loop subdivision
method, which generates new vertices by a weighted sum
of the existing vertices. Then, the depth of the dense 3D
model is modified with an average 3D depth-map in order
to preserve facial structure more realistically. Finally, all
249 3D models with expression changes are combined
to generate a 3DMM for a compact representation. The
first session of the Multi-PIE database, consisting of 249
persons with expression and illumination changes, is used
for the modeling. Unlike typical 3DMMs, our model can
generate 3D human faces more realistically and efficiently
(2-3 seconds on P4 machine) under diverse illumination
conditions.
link
Sunday, August 02, 2009
Lab Meeting August 3rd, 2009 (Jeff): Angular problem after loop closing
And try to propose some methods to slove this problem.
Thursday, July 30, 2009
MIT Talk: Biomedical Imaging and Analysis Seminar
Speaker: Arno Klein , Columbia University
Establishing correspondences across brains for the purposes of comparison and group analysis is almost universally done by registering images to one another either directly or via a template. However, there are many registration algorithms to choose from. The first part of this talk will give an overview of a recent evaluation study comparing fully automated nonlinear deformation methods applied to brain image registration (Klein et al. 2009). This study was restricted to volume-based methods, and an ongoing extension of this study is the first known to the authors that directly compares some of the most accurate of these methods with surface-based registration methods, as well as the first study to compare registrations of whole-head and de-skulled brain images. More than 6,000 registrations between 40 manually labeled brain images have been performed so far by the volume-based algorithms ART and SyN and the surface-based algorithms FreeSurfer and Spherical Demons. We used permutation tests and indifference-zone ranking to compare the overlap performance for eight scenarios: ART and SyN on brain images with and without skulls, SyN, FreeSurfer, and Spherical Demons via custom templates, and FreeSurfer via its default atlas.
Wednesday, July 29, 2009
Phd oral exam : Large Scale Scene Matching for Graphics and Vision
James H. Hays
11:00 AM, 7220 Wean Hall
Thesis Oral
Title: Large Scale Scene Matching for Graphics and Vision
Abstract:
Our visual experience is extraordinarily varied and complex. Thediversity of the visual world makes it difficult for computervision to understand images and for computer graphics tosynthesize visual content. But for all its richness, it turns outthat the space of "scenes" might not be astronomically large.With access to imagery on an Internet scale, regularities start toemerge - for most images, there exist numerous examples ofsemantically and structurally similar scenes. Is it possible tosample the space of scenes so densely that one can use similarscenes to "brute force" otherwise difficult image understandingand manipulation tasks? This thesis is focused on exploiting andrefining large scale scene matching to short circuit thetypical computer vision and graphics pipelines for imageunderstanding and manipulation.
First, in "Scene Completion" we patch up holes in images bycopying content from matching scenes. We find scenes so similarthat the manipulations are undetectable to naive viewers and wequantify our success rate with a perceptual study. Second, in"im2gps" we estimate geographic properties and globalgeolocation for photos using scene matching with a database of 6million geo-tagged Internet images. We geolocate sequences ofphotos four times as accurately as the single image case bymodelling the global spatiotemporal statistics of photographers.We introduce a range of features for scene matching and use them,together with lazy SVM learning, to dramatically improve scenematching -- doubling the performance of single image geolocationover our baseline method. Third, we study human photo geolocationto gain insights into the geolocation problem, our algorithms, andhuman scene understanding. This study shows that our algorithmssignificantly exceed human geolocation performance. Finally, weuse our geography estimates, as well as Internet text annotations,to provide context for deeper image understanding, such as objectdetection.
Thesis Committee:Alexei A. Efros, ChairMartial HebertJessica K. HodginsTakeo KanadeRichard Szeliski, Microsoft Research
Saturday, July 18, 2009
Lab Meeting August 3, 2009(Shao-Chen): Time-bounded Lattice for Efficient Planning in Dynamic Environments(ICRA09)
Saturday, July 11, 2009
Lab Meeting July 13, 2009 (Alan): An Embodied Cognition Approach to Mindreading Skills for Socially Intelligent Robots (IJRR 2009)
Authors: Cynthia Breazeal, Jesse Gray, Matt Berlin
Abstract
Future applications for personal robots motivate research into developing robots that are intelligent in their interactions with people. Toward this goal, in this paper we present an integrated socio-cognitive architecture to endow an anthropomorphic robot with the ability to infer mental states such as beliefs, intents, and desires from the observable behavior of its human partner. The design of our architecture is informed by recent findings from neuroscience and embodies cognition that reveals how living systems leverage their physical and cognitive embodiment through simulation-theoretic mechanisms to infer the mental states of others. We assess the robot’s mindreading skills on a suite of benchmark tasks where the robot interacts with a human partner in a cooperative scenario and a learning scenario. In addition, we have conducted human subjects experiments using the same task scenarios to assess human performance on these tasks and to compare the robot’s performance with that of people. In the process, our human subject studies also reveal some interesting insights into human behavior.
Link
Friday, July 10, 2009
Lab Meeting 7/13, 2009(Casey): A Closed-Form Solution to Non-Rigid Shape and Motion Recovery
Lab Meeting 7/13, 2009(fish60): Autonomous Robot Navigation in Outdoor Cluttered Pedestrian Walkways
Yoichi Morales, Alexander Carballo, Eijiro Takeuchi, Atsushi Aburadani, and Takashi Tsubouchi
Journal of Field Robotics 26(8), 609–635 (2009)
Abstract:
This paper describes an implementation of a mobile robot system for autonomous navigationin outdoor concurred walkways. The task was to navigate through nonmodified pedestrian paths with people and bicycles passing by. The proposed approach proved to be robust for outdoor navigationin cluttered and crowded walkways, first on campus paths and then running the challenge course multiple times between trials and the challenge final. The paper reports experimental results and overall performance of the system. Finally the lessons learned are discussed. The main contribution of this work is the report of a system integration approach for autonomous outdoor navigation and its evaluation.
Link
Thursday, June 25, 2009
News: Language may be key to theory of mind
Language may be key to theory of mind
- 15:55 23 June 2009 by Anil Ananthaswamy
- For similar stories, visit the The Human Brain Topic Guide
How blind and deaf people approach a cognitive test regarded as a milestone in human development has provided clues to how we deduce what others are thinking.
Understanding another person's perspective, and realising that it can differ from our own, is known as theory of mind. It underpins empathy, communication and the ability to deceive
– all of which we take for granted. Although our theory of mind is more developed than it is in other animals, we don't acquire it until around age four, and how it develops is a mystery.
See the full article.
Sunday, June 21, 2009
Lab Meeting June 21th, 2009 (swem): Improved Inverse-Depth Parameterization for Monocular Simultaneous Localization and Mapping
Simultaneous Localization and Mapping
In: ICRA2009
Author: E. Imre, M.-O. Berger, N. Noury
Abstract:
Inverse-depth parameterization can successfully
deal with the feature initialization problem in monocular
simultaneous localization and mapping applications. However,
it is redundant, and when multiple landmarks are initialized
from the same image, it fails to enforce the “common origin”
constraint. The authors propose two new variants that
addresses both of these issues. The experimental results indicate
that the proposed approach achieves a better performance at a
lower computational cost.
[ Link ]
Saturday, June 20, 2009
Intelligence Seminar: Action Perception, June 23, 2009
June 23, 2009 (note special place)
3:30 pm
NSH 1507
Host: Jaime Carbonell
For meetings, contact Michelle Pagnani (pagnani@cs.cmu.edu).
Action Perception
Robert Thibadeau
Seagate Research
Abstract:
The human perception of actions has barely been studied, but this study of action perception promises to provide a wealth of interesting hypotheses regarding cognitive processing. Action perception is distinct from motion perception in that the direct perception of causation is central to the percept. Among the interesting hypotheses is that it can be hypothesized that what we know as thought and reasoning is where we perceive and plan actions. Another hypothesis is that what we know as logic and mathematics derives from our direct perceptions of causation in the actions we perceive and think about.
I will present a study that attempts to estimate the scale of computation needed to implement a system for visually perceiving meaningful actions and non-trivially producing an English narration of what is being visually perceived, as well as answering questions about what is visually perceived. The scale of the computation for learning could easily reach exaflops over distributed datasets (HADOOP or MapReduce style).
This study is partly based on my work (Thibadeau, 1986), and Doug Rohde's 2002 dissertation (http://tedlab.mit.edu:16080/~
(Simon and Rescher 1966 From Wikipedia, Causality)
Derivation theories
The Nobel Prize holder Herbert Simon and Philosopher Nicholas Rescher[20] claim that the asymmetry of the causal relation is unrelated to the asymmetry of any mode of implication that
contraposes. Rather, a causal relation is not a relation between values of variables, but a function of one variable (the cause) on to another (the effect). So, given a system of equations, and a set of
variables appearing in these equations, we can introduce an asymmetricrelation among individual equations and variables that corresponds perfectly to our commonsense notion of a causal ordering. The system of equations must have certain properties, most importantly, if some values are chosen arbitrarily, the remaining values will be determined uniquely through a path of serial discovery that is perfectly causal. They postulate the inherent serialization of such a system of equations may correctly capture causation in all empirical fields, including physics and economics.
Sunday, June 14, 2009
RSS 2009 paper: Non-parametric Learning To Aid Path Planning Over Slopes
Sisir Karumanchi, Thomas Allen, Tim Bailey and Steve Scheding
ARC Centre of Excellence For Autonomous Systems (CAS),
Australian Centre For Field Robotics (ACFR),
The University of Sydney,
NSW. 2006, Australia.
Abstract—This paper addresses the problem of closing the loop from perception to action selection for unmanned ground vehicles, with a focus on navigating slopes. A new non-parametric learning technique is presented to generate a mobility representation where maximum feasible speed is used as a criterion to classify the world. The inputs to the algorithm are terrain gradients derived from an elevation map and past observations of wheel slip. It is argued that such a representation can aid in path planning with improved selection of vehicle heading and operating velocity in off-road slopes. Results of mobility map generation and its benefits to path planning are shown.
Link
Saturday, June 13, 2009
RSS 2009 Paper: Generalized-ICP
RSS 2009 Paper: Large Scale Graph-based SLAM using Aerial Images as Prior Information
Lab Meeting June 15th, 2009 (Andi): Progress Report
Using the Distribution Theory to Simultaneously Calibrate the Sensors of a Mobile Robot
Author: Agostino Martinelli
Abstract:
This paper introduces a simple and very efficient strategy to extrinsically calibrate a bearing sensor (e.g. a camera) mounted on a mobile robot and simultaneously estimate the parameters describing the systematic error of the robot odometry system. The paper provides two contributions. The first one is the analytical computation to derive the part of the system which
is observable when the robot accomplishes circular trejectories. This computation consists in performing a local decomposition of the system, based on the theory of distributions. In this respect, this paper represents the first application of the distribution theory in the frame-work of mobile robotics. Then, starting from this decomposition, a method to efficiently estimate the parameters describing both the extrinsic bearing sensor calibration and the odometry calibration is derived (second contribution). Simulations and experiments with the robot e-Puck equipped
with encoder sensors and a camera validate the approach.
Link:
RSS05
http://www.roboticsproceedings.org/rss05/p11.pdf
http://hal.inria.fr/docs/00/35/30/79/PDF/RR-6796.pdf
Thursday, June 11, 2009
Sunday, June 07, 2009
Lab Meeting June 8th, 2009(ZhenYu):Vertical Line Matching for Omnidirectional Stereovision Images
Authors: Guillaume Caron and El Mustapha Mouaddib
[Link]
Saturday, June 06, 2009
Lab Meeting June 8th, 2009 (Any): CRF-Filters
Sunday, May 31, 2009
Some talks at ICRA 2009.
Tuesday, May 26, 2009
ICRA2009 Awards
Moving Obstacle Detection in Highly Dynamic Scenes [Local Copy][Attachment]
Best Automation Paper and Best Student Paper
Design and Calibration of a Microfabricated 6-Axis Force-Torque Sensor for Microrobotic Applications [Local Copy]
Best Conference Paper
Towards a Navigation System for Autonomous Indoor Flying [Local Copy]
Monday, May 25, 2009
Lab Meeting June 1st, 2009 (Jeff): Modeling RFID Signal Strength and Tag Detection for Localization and Mapping
Authors: Dominik Joho, Christian Plagemann and Wolfram Burgard
Abstract:
In recent years, there has been an increasing interest within the robotics community in investigating whether Radio Frequency Identification (RFID) technology can be utilized to solve localization and mapping problems in the context of mobile robots. We present a novel sensor model which can be utilized for localizing RFID tags and for tracking a mobile agent moving through an RFID-equipped environment. The proposed probabilistic sensor model characterizes the received signal strength indication (RSSI) information as well as the tag detection events to achieve a higher modeling accuracy compared to state-of-the-art models which deal with one of these aspects only. We furthermore propose a method that is able to bootstrap such a sensor model in a fully unsupervised fashion. Real-world experiments demonstrate the effectiveness of our approach also in comparison to existing techniques.
Link:
ICRA2009
http://www.informatik.uni-freiburg.de/~joho/publications/joho09icra.html
Sunday, May 17, 2009
Lab Meeting May 25, 2009 (fish60):
Brenna Argall, Sonia Chernova and Manuela Veloso. A Survey of Robot Learning from Demonstration. Robotics and Autonomous Systems. Vol. 57, No. 5, pages 469-483, 2009.
Link
abstract:
We present a comprehensive survey of robot Learning from Demonstration (LfD), a technique that develops policies from example state to action mappings. We introduce the LfD design choices in terms of demonstrator, problem space, policy derivation and performance, and contribute the foundations for a structure in which to categorize LfD research. Specifically, we analyze and categorize the multiple ways in which examples are gathered, ranging from teleoperation to imitation, as well as the various techniques for policy derivation, including matching functions, dynamics models and plans. To conclude we discuss LfD limitations and related promising areas for future research.
Wednesday, May 13, 2009
MIT CSAIL Technical Report: Scene Classification with a Biologically Inspired Method
Monday, May 11, 2009
VASC Seminar: A Data-Driven Vision Compiler for Automatic Object Pose Recognition
Monday, May 11, 2009
3:30p-4:30p
NSH 1507
A Data-Driven Vision Compiler for Automatic Object Pose Recognition
Rosen Diankov
Carnegie Mellon University, Robotics Institute
Abstract:
This presentation focuses on an object-specific vision system that detects and extracts the precise 6D pose of objects in an image. The system builds a data-driven statistical model of the expected features of an object's surface and combines this with a discrete search method
to extract the pose of all object. The training phase of the vision system can be interpreted as a compiler that automatically analyzes the statistics of how the features are distributed on the object and determines a feature set's stability and discriminable power. This compilation phase requires the precise CAD model of an object along with a training set of real-world images. After compilation, a CAD-independent model of how features relate with respect to theobject's pose and inter-relate with each other is created. These relationships allow both point-based features like SIFT and edge-based features to be used simultaneously when computing the 6D pose of an
object. Using this data-driven model, we employ a discrete randomized search with RANSAC to find the poses of all instances of the object in a novel image.
Bio:
Rosen Diankov graduated from University of California Berkeley in 2006 with Electrical Engineering and Computer Science, and Applied Math degrees. At the moment he is a PhD graduate student at the Robotics Institute at Carnegie Mellon University. Rosen's main research focus is tackling the robotics problem: combining perception, planning, and control into one coherent framework. Up until now he has worked on several vision and planning systems involving autonomous robots in everyday scenarios both in the United States and Japan.
VASC Seminars are sponsored by Tandent Vision Science, Inc.
Lab Meeting May 25, 2009(Chung-Han) COLD : The CoSy Localization Database
Author : A. Pronobis, B. Caputo
Abstract : Two key competencies for mobile robotic systems are localizationand semantic context interpretation. Recently, vision has become themodality of choice for these problems as it provides richer and moredescriptive sensory input. At the same time, designing and testingvision-based algorithms still remains a challenge, as large amounts ofcarefully selected data are required to address the high variability ofvisual information. In this paper we present a freely available databasewhich provides a large-scale, flexible testing environment forvision-based topological localization and semantic knowledge extractionin robotic systems. The database contains 76 image sequencesacquired in three different indoor environments across Europe. Acquisitionwas performed with the same perspective and omnidirectionalcamera setup, in rooms of different functionality and under variousconditions. The database is an ideal testbed for evaluating algorithmsin real-world scenarios with respect to both dynamic and categoricalvariations.
[Full text]
Tuesday, May 05, 2009
CMU talk: Fast Feature Detection and Stochastic Parameter Estimation of Road Shape using Multiple LIDAR
Fast Feature Detection and Stochastic Parameter Estimation of Road Shape using Multiple LIDAR
Kevin Peterson
PhD Student, Robotics Institute, Carnegie Mellon University
Thursday, May 7th, 2009
Developers of autonomous vehicles must overcome significant challenges before these vehicles can operate around human drivers. In urban environments autonomous cars will be required to follow complex traffic rules regarding merging and queuing, navigate in close proximity to other drivers, and safely avoid collision with pedestrians and fixed obstacles near the road. Knowledge of the location and shape of the roadway near then autonomous car is fundamental to these behaviors. While it is tempting to build an /a priori /GPS-registered map of the road network, the possibility for change in the road structure (e.g. new roads, construction, etc.) precludes the use of maps alone. It is therefore necessary to detect and track roads in real-time.
A rich body of work exists in the area of road tracking. Although some early work was performed on unimproved roads, a majority of the available research focuses on paved roads and highways. Additionally, a vast majority of this work focuses on the use of cameras as the single sensing modality. In this talk I will present a framework for road tracking that uses a particle filter to fuse several sources of data including LIDAR and video. The approach enables road tracking on unimproved roads and, because of the flexible nature of the particle filter, can easily be extended to incorporate new forms of data. I will present results from our preparations for the DARPA Urban Challenge.
Speaker Bio: Kevin Peterson’s research focuses on perception techniques for robust autonomous vehicles. Kevin holds a B.S. and M.S. in Electrical and Computer Engineering from Carnegie Mellon University in Pittsburgh, Pennsylvania and is currently pursuing a PhD in Robotics also from Carnegie Mellon University. Kevin has built software systems for many autonomous systems with applications ranging from cave exploration to unexploded ordinance cleanup. Notably, Kevin led Red Team Too, one of CMUs entries in the 2005 DARPA Grand Challenge and was a participant on Tartan Racing, CMUs winning entry in the 2007 Urban Grand Challenge. Since then, he has been applying inverse optimal control techniques to build models of pedestrian motion in structured and unstructured environments.
Saturday, May 02, 2009
NTU PhD oral: Mobile Agent Enhanced Service-Oriented Smart Home Architecture and its Human-System Interaction Framework and Algorithm
Mobile Agent Enhanced Service-Oriented Smart Home Architecture and its Human-System Interaction Framework and Algorithm
Chao-Lin Wu
Date: May 7, 2009 (Thursday)
Time: 10 am~ 12 noon
Place: CSIE R340
Advisor: Li-Chen Fu
NTU PhD oral: Computer Vision Techniques for Effective Pedestrian Detection
Yu-Ting Chen
Date: May 6, 2009 (Wednesday)
Time: 10 am~12 noon
Place: CSIE R440
Advisors: Chu-Song Chen and Yi-Ping Hung
口試委員:
(1) 校內:洪一平老師、陳祝嵩老師、傅立成老師、王傑智老師
(2) 校外:陳世旺老師、王聖智老師、賴尚宏老師、鍾國亮老師
MIT PhD Thesis: Robust and Efficient Robotic Mapping
Title: Robust and Efficient Robotic Mapping
Author: Edwin B. Olson
Date: June 2008
Abstract
Mobile robots are dependent upon a model of the environment for many of their basic functions. Locally accurate maps are critical to collision avoidance, while large-scale maps (accurate both metrically and topologically) are necessary for efficient route planning. Solutions to these problems have immediate and important applications to autonomous vehicles, precision surveying, and domestic robots.
Building accurate maps can be cast as an optimization problem: find the map that is most probable given the set of observations of the environment. However, the problem rapidly becomes difficult when dealing with large maps or large numbers of observations. Sensor noise and non-linearities make the problem even more difficult— especially when using inexpensive (and therefore preferable) sensors.
This thesis describes an optimization algorithm that can rapidly estimate the maximum likelihood map given a set of observations. The algorithm, which iteratively reduces map error by considering a single observation at a time, scales well to large environments with many observations. The approach is particularly robust to noise and non-linearities, quickly escaping local minima that trap current methods. Both batch and online versions of the algorithm are described.
In order to build a map, however, a robot must first be able to recognize places that it has previously seen. Limitations in sensor processing algorithms, coupled with environmental ambiguity, make this difficult. Incorrect place recognitions can rapidly lead to divergence of the map. This thesis describes a place recognition algorithm that can robustly handle ambiguous data.
We evaluate these algorithms on a number of challenging datasets and provide quantitative comparisons to other state-of-the-art methods, illustrating the advantages of our methods.
Link: pdf
CMU talk: Click Chain Model in Web Search
Speaker: Fan Guo
Date: Monday, May 4, 2009
Title: Click Chain Model in Web Search
Abstract:
Given a terabyte click log, can we build an efficient and effective click model? It is commonly believed that web search click logs are a gold mine for search business, because they reflect users' preference over web documents presented by the search engine. Click models provide a principled approach to inferring user-perceived relevance of web documents, which can be leveraged in numerous applications in search businesses. Due to the huge volume of click data, scalability is a must. I will present the click chain model, which is based on a solid, Bayesian framework. It is both scalable and incremental, perfectly meeting the computational challenges imposed by the voluminous click logs that constantly grow.
Joint work with Chao Liu, Anitha Kannan, Tom Minka, Michael Taylor, Yi-Min Wang and Christos Faloutsos.
Friday, May 01, 2009
NTU talk: AHCBoost: Boosting for Multi-class Classification
Speaker: Prof. Andy Chen-Hai Tsao, National Dong Hwa University
Time: 02:20pm, May 22 (Friday), 2009
Place: Room 102, CSIE building
Abstract:
AdaBoost is one of the important ensemble classifiers developed in the last decade. However, there are difficulties in applying AdaBoost for multi-class classifications. In this study, we introduce the adjustable hyperbolic cosine loss to develop a new boosting algorithm, AHCBoost. Our experiments on benchmark data sets suggest that AHCBoost is very competitive with or even better than the multi-class classifiers SVM and glmBoost. In addition to the fast reduction of training and testing errors, AHCBoost is relatively immune to overfitting and requires little parameter tuning. Some experiments for exploring the potential and limitation of using AHCBoost for ordinal response will also be reported.
Short Biography:
Professional Positions
Department of Applied Math, National Dong Hwa University
* Professor: 2005 to present
* Associate Professor: 1997 to 2005
* Assistant Professor : 1995 to 1997
Institute of Math Statistics, National Chung Cheng University
* Visiting Associate Professor: 1994 to 1995
Thursday, April 30, 2009
CMU talk: Learning to Search: Structured Prediction Techniques for Imitation Learning
Learning to Search: Structured Prediction Techniques for Imitation Learning
Nathan D. Ratliff
Carnegie Mellon University
May 01, 2009
Abstract: Modern robots successfully manipulate objects, navigate rugged terrain, drive in urban settings, and play world-class chess. Unfortunately, programming these robots is challenging, time-consuming and expensive; the parameters governing their behavior are often unintuitive, even when the desired behavior is clear and easily demonstrated. Inspired by successful end-to-end learning systems such as neural network controlled driving platforms (Pomerleau, 1989), learning-based "programming by demonstration" has gained currency as a method to achieve intelligent robot behavior. Unfortunately, with highly structured algorithms at their core, it is not clear how to effectively and efficiently train modern robotic systems using classical learning techniques. Rather than redefining robot architectures to accommodate existing learning algorithms, in this thesis I develop learning techniques that leverage the performance of modern robotic components.
My presentation begins with a discussion of a novel imitation learning framework we call Maximum Margin Planning which automates finding a cost function for optimal planning and control algorithms such as A*. In the linear setting, this framework has firm theoretical backing in the form of strong generalization and regret bounds. Further, I have developed practical nonlinear generalizations that are effective and efficient for real-world problems. This framework reduces imitation learning to a modern form of machine learning known as Maximum Margin Structured Classification (Taskar et al. 2005); these algorithms, therefore, apply both specifically to training existing state-of-the-art planners, as well as broadly to solving a range of structured prediction problems of importance in learning and robotics.
In difficult high-dimensional planning domains, such as those found in many manipulation problems, high-performance planning technology remains a topic of much research. I will present some recent work which moves toward simultaneously advancing this technology while retaining the learnability developed above.
I'll demonstrate our algorithms on a range of applications including overhead navigation, quadrupedal locomotion, heuristic learning, manipulation planning, grasp prediction, driver prediction, pedestrian prediction, optical character recognition, and LADAR classification.
Saturday, April 25, 2009
CMU talk: Camera and LIDAR Fusion for Mapping in Dark Environments
Camera and LIDAR Fusion for Mapping in Dark Environments
Uland Wong
PhD Student, Robotics Institute, Carnegie Mellon University
Thursday, April 30th, 2009
Abstract: Unlit diffuse environments like subterranean voids on earth and planetary surfaces elsewhere are of great interest for robotic exploration and exploitation. These environments pose unique obstacles and constraints, including the necessity for active perception and illumination. However, uniformity of albedo, lack of external lighting and known surface reflectance provide additional assumptions which can be used to enhance 3D-mapping and photographic data collected from robots. This talk presents a method for improving the accuracy of super-resolution point clouds by fusing actively illuminated HDR camera imagery with LIDAR data in dark Lambertian environments. The key approach is shape recovery from estimation of the illumination function and integration in a Markov Random Field (MRF) framework. Experimental results collected from a virtual reconstruction of the Bruceton Research Mine in Pittsburgh, PA are also presented.
Tuesday, April 21, 2009
NTU talk: Community Discovery in Dynamic, Rich Media Social Networks
Speaker: Yu-Ru Lin, PhD Candidate, Arizona State University.
Time: 3:50pm ~ 5:00pm, Wednesday, March 22, 2009.
Place: Room 111, CSIE building
Abstract: With the rapid proliferation of different types of social media, such as instant messaging (e.g., AIM, MSN, Skype), media sharing sites (e.g., Flickr, YouTube), blogs (e.g., Blogger, WordPress, LiveJournal), wikis (e.g., Wikipedia, PBWiki), microblogs (e.g., Twitter, Jaiku), social networks (e.g., MySpace, Facebook), to mention a few, users routinely produce (e.g. blogs) and consume media (e.g. YouTube) as well as interact with each other through a wide array of functionality provided by various social media. These social media depend largely on implicit communities of online users to deliver value. Identifying and analyzing the dynamics of such latent communities can lead to improved functionality of the social media as well as provide insight into the design of future online collaborative services. The problem is particularly important in the enterprise domain where extracting emergent community structure on enterprise social media, can help in forming new collaborative teams, in expertise discovery, and guide long term enterprise reorganization.
In this talk, I will cover three aspects of community analysis in dynamic, rich media social networks: (1) Community evolution – How do we identify communities in large scale, dynamic social networks, and analyze their structures and evolutions? I will introduce a robust unified approach that discovers communities and captures their evolution with temporal smoothness given by historic community structure. (2) Community summarization – How do we summarize community activities, in order to trace community interests or retrieve community generated content? I will present a summarization framework that characterizes the time-evolving patterns of social activities with associated media objects in a community. (3) Multi-relational communities – How do we discover communities when the social networks exist in a highly connected web of contexts (e.g., social groups, geographic locations, time, etc.)? I will discuss a novel multi-relational non-negative tensor decomposition algorithm that aims to solve this problem. I will also show the effectiveness of these techniques in real world datasets collected from the blogosphere, an enterprise, Flickr, Digg, etc.
Short Biography: Yu-Ru Lin is currently a Ph.D. student in the School of Computing and Informatics at Arizona State University, with a concentration in Arts, Media and Engineering. Her advisor is Dr. Hari Sundaram. Her research interests include problems relating to dynamic multi-relational social network analysis – in particular, community dynamics, social information summarization and representation. Her research focuses on extracting human communities that collaborate around certain topics or media sharing activities. She has proposed non-negative matrix/tensor factorization techniques for analyzing community structures and evolutions in online social networks, as well as time-varying social relational data. Her work has been published in leading international conferences and journals. (Her publication can be found at http://www.public.asu.edu/~ylin56/pub.html.)
She has worked at NEC Labs America and IBM TJ Watson Research Center as a summer intern in 2006, 2007 and 2008. She has received awards including AME Student Excellence Award (2007 and 2008) and IBM PhD Fellowship Award (2009). She holds an M.S. and B.S. degree in Computer Science from National Chiao Tung University, Taiwan.
NTU talk: Mining Geotagged Photos for Semantic Understanding
Speaker: Dr. Jiebo Luo, IEEE Fellow, Senior Principal Scientist with the Kodak Research Laboratories.
Time: 2:30pm ~ 3:40pm, Wednesday, March 22, 2009.
Place: Room 111, CSIE building
Abstract:
Semantic understanding based only on vision cues has been a challenging problem. This problem is particularly acute when the application domain is unconstrained photos available on the Internet or in personal repositories. In recent years, it has been shown that metadata captured with pictures can provide valuable contextual cues complementary to the image content and can be used to improve classification performance. With the recent geotagging phenomenon, an important piece of metadata available with many geotagged pictures is GPS information. We will describe a number of novel ways to mine GPS information in a powerful contextual inference framework that boosts the accuracy of semantic understanding. With integrated GPS-capable cameras on the horizon and geotagging on the rise, this line of research will revolutionize event recognition and media annotation.
Short Biography: Jiebo Luo is a Senior Principal Scientist with the Kodak Research Laboratories in Rochester, NY. He received a B.S. degree and M.S. degree in Electrical Engineering from the University of Science and Technology of China (USTC) in 1989 and 1992, respectively, and a Ph.D. degree in Electrical Engineering from the University of Rochester in 1995. His research interests include signal and image processing, pattern recognition, computer vision, and the related multi-disciplines such as multimedia data mining, biomedical informatics, computational photography, and human-computer interaction. Dr. Luo has authored over 130 technical papers and holds 50 granted US patents. Dr. Luo actively participates in numerous technical conferences, including serving as the chair of the 2008 ACM International Conference on Content-based Image and Video Retrieval (CIVR), an area chair of the 2008 IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), a program co-chair of the 2007 SPIE International Symposium on Visual Communication and Image Processing (VCIP), a member of the Organizing Committee of the 2008 ACM Multimedia Conference, 2006 & 2008 IEEE International Conference on Multimedia and Expo (ICME) and 2002 IEEE International Conference on Image Processing (ICIP), and the chair of the IEEE CVPR Workshop on Semantic Learning Application in Multimedia (SLAM) since its inception in 2006. He is the Editor-in-Chief for the Journal of Multimedia (Academy Publisher). Currently, he is also on the editorial boards of the IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), the IEEE Transactions on Multimedia (TMM), Pattern Recognition (PR), and Journal of Electronic Imaging (JEI). He is a guest editor for a number of influential special issues, including “Image Understanding for Digital Photos” (PR, 2005), “Real-World Image Annotation and Retrieval” (TPAMI, 2008), “Event Analysis in Video” (IEEE Transactions on Circuits and Systems for Video Technology, 2008), “Integration of Content and Context for Multimedia Management” (TMM, 2009), and “Probabilistic Graphic Models in Computer Vision” (TPAMI, 2009). Dr. Luo is an adjunct professor at Rochester Institute of Technology, as well as the co-dvisor or thesis committee member of many PhD and MS graduate students in various US universities. He is a Kodak Distinguished, a Fellow of SPIE for achievements in electronic imaging and visual communication, and a Fellow of IEEE for contributions to semantic image understanding and intelligent image processing.
Saturday, April 18, 2009
Lab Meeting April 20, 2009 (Casey): Avoiding moving outliers in visual SLAM by tracking moving objects
Authors: Somkiat Wangsiripitak and David W Murray
Lab Meeting April 20, 2009 (Yuchun): An Affective Guide Robot in a Shopping Mall
Authors: Takayuki Kanda, Masahiro Shiomi, Zenta Miyashita, Hiroshi Ishiguro and Norihiro Hagita HRI 2009: 173-180
Abstract:
To explore possible robot tasks in daily life, we developed a guide robot for a shopping mall and conducted a field trial with it. The robot was designed to interact naturally with customers and to affectively provide shopping information. It was also designed to repeatedly interact with people to build a rapport; since a shopping mall is a place people repeatedly visit, it provides the chance to explicitly design a robot for multiple interactions. For this capability, we used RFID tags for person identification. The robot was semi- autonomous, partially controlled by a human operator, to cope with the difficulty of speech recognition in a real environment and to handle unexpected situations.
A field trial was conducted at a shopping mall for 25 days to observe how the robot performed this task and how people interacted with it. The robot interacted with approximately 100 groups of customers each day. We invited customers to sign up for RFID tags and those who participated answered questionnaires. The results revealed that 63 out of 235 people in fact went shopping based on the information provided by the robot. The experimental results suggest promising potential for robots working in shopping malls.
[pdf]
Thursday, April 16, 2009
CMU talk: Looking without Seeing is in fact Seeing without Knowing -- Insights from Gaze-tracked Change Blindness Studies
Looking without Seeing is in fact Seeing without Knowing
-- Insights from Gaze-tracked Change Blindness Studies
Stella X. Yu
Clare Boothe Luce Assistant Professor
Computer Science @ Boston College
http://www.cs.bc.edu/~syu
Tuesday, April 28
Abstract: Change blindness experiments demonstrate that human vision often neglects certain aspects of the visual scene while attending to others. Using a gaze-tracked flicker paradigm and synthetic images equally rendered in three fundamental features, we explore whether and how innate feature processing might be responsible for the blindness to a change. Our analysis of detection accuracy, detection time, and gaze patterns in this active visual search task reveals distinctive feature extraction, discrimination, and selectivity of size, color, and orientation that underly different behaviours of change blindness. With an array of two-feature stimuli where a single element could change in either feature dimension, we discover that what is changing is sensed long before the subject consciously detects the change, and the change detection task is not accomplished in a single thread of searching for a nonspecific change, but in three separate threads: sensing what the change is, localizing where it might be, and discerning how it is actualized.
Bio: Stella X. Yu got her Ph.D. from the School of Computer Science at Carnegie Mellon University, where she studied robotics at the Robotics Institute and vision science at the Center for the Neural Basis of Cognition. She then went to the University of California, Berkeley to continue her research on computer vision. Her research interests are visual perception and computer vision. Since she joined Boston College, Dr. Yu has been developing an interdisciplinary curriculum and research agenda around art and vision. Her 5-year NSF CAREER proposal, entitled Art and Vision: Scene Layout from Pictorial Cues, was awarded in 2007. Her recent works include image segmentation, object matching, spatial layout categorization and inference, change blindness, and brightness perception.
Tuesday, April 14, 2009
CMU talk: Multi-robot Coordination for Domains with Intra-path Constraints
Multi-robot Coordination for Domains with Intra-path Constraints
E. Gil Jones
PhD Candidate
CMU Robotics Institute
Thursday, April 16nd, 2009
Abstract
Many applications require teams of robots to cooperatively execute complex tasks. Among these domains are those where successful coordination solutions must respect constraints that occur on the intra-path level. This work focuses on multi-agent coordination for disaster response with intra-path constraints, a compelling application that is not well addressed by current coordination methods. In this domain a group of fire trucks agents attempt to address a number of fires that are occurring throughout a city in the wake of a large-scale disaster. The disaster has not only caused fires but has also caused many city roads to be blocked by debris, making them impassable; bulldozer robots also operating in the domain can clear the debris. The coordination solution must determine not only a task allocation but also what routes the fire trucks should take given the intra-path precedence constraints and which bulldozers should be assigned to satisfy those constraints. This talk will focus on two main techniques for determining multi-robot coordination solutions for domains with intra-path constraints. The first technique uses tiered auctions, a novel market-based method. The second technique uses centralized genetic algorithms. The approaches are compared in terms of solution quality and computation time in a simulated disaster response domain.
Bio
Gil is a fifth year Ph.D. student at the Robotics Institute, and is co-advised by Bernardine Dias and Tony Stentz. He received his BA in Computer Science from Swarthmore College in 2001, and spent two years as a software engineer at Bluefin Robotics - manufacturer of autonomous underwater vehicles - in Cambridge, Mass.
Saturday, April 11, 2009
CMU talk: Inferring Object Attributes
Inferring Object Attributes
Derek Hoiem
Assistant Professor, University of Illinois at Urbana Champaign
April 10, 2009
Abstract: Ultimately, the goal of computer vision is to make useful inferences from imagery, and a big part of that is knowing something about the properties of nearby objects. In this talk, I'll describe our recent work on learning to identify object attributes, such as parts, materials, or shape, from images in a way that generalizes to new object categories. The tricky part is training classifiers that really predict the intended attribute, and not ones that are correlated through familiar object categories. Once we can predict attributes, we can say what is unusual about an object and more easily learn to recognize new objects Sometimes we can even recognize new object categories from a purely verbal description (e.g., a goat has four legs, horns, and is furry).
This work is with Ali Farhadi, Ian Endres, and David Forsyth at UIUC.
Speaker Bio.: Derek Hoiem is a new assistant professor at University of Illinois at Urbana Champaign. Derek researches object recognition, segmentation, 3d reconstruction from images, and other aspects of computer vision that are related to scene understanding. He recently (2007) graduated from the Robotics Institute under the tutelage of Alyosha Efros and Martial Hebert and looks forward to visiting. By request, Derek will share a little of his perspective in transitioning from being a grad student at CMU to a professor at UIUC.
CMU talk: Fundamental Limits of Imaging in Scattering Media
Monday, April 13, 2009
Fundamental Limits of Imaging in Scattering Media
Tali Treibitz
Technion
Abstract:
Scattering media exist in bad weather, liquids, biological tissue and even solids. Images taken into scattering media suffer from resolution loss photometric as well as geometric problems. In addition, high noise levels impose resolution limits, even if there is no blur. These problems inhibit vision in such media. In the talk I give an overview of our contributions in this subject:
- Resolution limits imposed by noise
- Limits in polarization-based dehazing
- Geometry limits: The non-single viewpoint nature of imaging systems looking into water through a flat glass.
Bio:
Tali Treibitz received her BA degree in computer science from the Technion-Israel Institute of Technology in 2001. She is currently a Ph.D. candidate in the department of Electrical Engineering, Technion. Her research involves physics-based computer vision. She is also an active PADI open water scuba instructor.
CMU talk: Discourse Structure from Topic Models in Text and Video
Speaker: Jacob Eisenstein (UIUC)
Venue: NSH 1507
Date: Monday, April 13, 2009
Title:
Discourse Structure from Topic Models in Text and Video
Abstract:
This talk describes how latent topic models can be used to discover discourse structure in unannotated text and video. Linguists have long believed that discourse topic shifts are marked by changes in the distribution of lexical items. This idea is called lexical cohesion, and can be formalized in a latent topic model, yielding substantial performance gains over previous heuristic approaches. More importantly, this Bayesian setting permits several interesting extensions:
(1) Explicit cue phrases for topic transitions are clearly relevant for segmentation, but can not be handled by previous unsupervised methods. I'll show how such cue phrases can be discovered without annotation and incorporated to improve segmentation.
(2) Topic segmentation can be applied to multimedia data by searching for self-similarity in visual communication. I'll present a topic segmenter for conversational speech that integrates lexical and gestural cohesion.
(3) Hierarchical structure can be discovered by modeling lexical cohesion as a multi-scale phenomenon, in which some words are governed by low-level subtopics, and others by the high-level topics. Inference is performed jointly across scale-levels, improving on greedy top-down approaches.
Overall, these extensions outperform the state-of-the-art on several tasks, and point the way to more comprehensive analysis of discourse structure through hierarchical Bayesian models.
BIO: Jacob Eisenstein is a Beckman Postdoctoral Fellow at the University of Illinois. He completed his doctorate at MIT in 2008 under the supervision of Regina Barzilay and Randall Davis. His thesis, titled "Gesture in Automatic Discourse Processing," won the 2008 George M. Sprowls award for best Doctoral theses in Computer Science at MIT. Working in the domain of computational linguistics, Jacob's research focuses on applying state-of-the-art structured learning techniques to discourse processing and visual communication.
Friday, April 10, 2009
Lab Meeting April 13,2009 (Jimmy): Mutual Localization in a Team of Autonomous Robots using Acoustic Robot Detection
Lab Meeting April 13,2009 (Gary) Subtle facial expression recognition using motion magnification
Authors: Sungsoo Park, Daijin Kim
Pattern Recognition Letters
Volume 30, Issue 7, 1 May 2009, Pages 708-716
Abstract
This paper proposes a novel method for subtle facial expression recognition that uses motion magnification to transform subtle expressions into corresponding exaggerated ones. Motion magnification consists of four steps: First, active appearance model (AAM) fitting extracts 70 facial feature points in the face image sequence. Second, the face image sequence is aligned using the three feature points (two eyes and nose tip). Third, the motion vectors of 27 feature points are estimated using the feature point tracking method. Finally, exaggerated facial expressions are obtained by magnifying the motion vectors of the 27 feature points. After motion magnification, the exaggerated facial expressions are recognized as follows: first, the shape and appearance features are obtained by projecting the exaggerated facial expression image to the AAM shape and appearance model. Second, support vector machines (SVM) are used to classify shape and appearance features. Experimental results show that proposed subtle facial recognition rate is 88.125% for the 80 facial expression images in the SFED2007 database.
link
Saturday, April 04, 2009
NTU talk: Skill Learning for Humanoid Robots
Time: 2:20 PM, May 8, 2009
Place: NTU CSIE
Hsien-I Lin
School of Electrical and Computer Engineering
Purdue University
West Lafayette, IN 47907-2035
sofin@purdue.edu
Abstract:
Recent advances in human-centered robots such as humanoid robots are driven by the projection that these robots will have a place in our society and in our daily life activities as assistive robots. Endowing these humanoid robots with the ability of skill learning will enable them to be versatile and skillful in performing various tasks. The problem of transferring human skills to humanoid robots raises tremendous research interest in studying human and robot motor skills. Our current research aims at developing a quantitative measure of robot motor capability of a humanoid root motor system for the application of transferring human skills to a humanoid robot.
We propose to employ an information-theory-based method to quantitatively represent the robot motor capability by a pseudo index of motor performance. This pseudo index of motor performance is derived from kinematics, dynamics, and control with the speed-accuracy constraint taken into consideration. With the speed-accuracy constraint, we are able to optimize the motor performance of a robot to accomplish a task by satisfying the task spatial and temporal constraints. Computer simulations and experimental work were performed on a 6 DOF PUMA robot to validate the performance of the proposed approach in measuring the robot motor capability of a root motor system.
Bio: Hsien-I Lin received the B.S. and M.S. degrees in Electrical and Control Engineering from National Chiao Tung University in 1997 and 1999, respectively, and he is currently a Ph.D. candidate in Electrical and Computer Engineering at Purdue University, West Lafayette, Indiana. Before beginning his academic career, he worked for the VIA technologies, Inc., Taipei, Taiwan during 2001-2003, after which he was a research assistant of the Department of Bio-Industrial Mechatronics Engineering at National Taiwan University during 2003-2004. Since then, he has been a research assistant of the Department of Electrical and Computer Engineering at Purdue University. His research interests are in the areas of human-robot interaction with emphasis on robot skill learning, intelligent systems, and neuro-fuzzy networks.
CMU talk: Fourier Theoretic Probabilistic Inference over Permutations
Friday, April 03, 2009
News: A Big-Screen Display as Mobile as Your Phone
Projector phones give us the first glimpse at a future where device size and display size are independent. The companies behind liquid crystal on silicon, digital light processing, and laser projection technologies each think they have a superior technology, but watch the video and decide for yourself.
View Now!
I just can not forget our previous work on the camera-projector systems. -Bob
Thursday, April 02, 2009
CMU talk: Next Generation Map Making
NAVTEQ
Monday, April 6, 2009
NAVTEQ is a leading global provider of digital map data. NAVTEQ maps drive most in-vehicle navigation systems, the top routing web sites, and the leading brands of wireless navigation devices. NAVTEQ continues to enhance the technologies used for collecting, analyzing, and delivering new content to a wide range of users and devices. Dr. Chen will discuss NAVTEQ's perspective on the hardware and software systems required to automatically create and maintain a navigable map through the use of high-end mobile data collection sensors and computer vision techniques. Dr. Chen will also present numerous research efforts based on video and LIDAR data collection as well as various challenging problems related to automatic feature extraction for mapping and navigation.
Bio:
Dr. Xin Chen currently works as a senior researcher in the Research and Emerging Technologies Department of NAVTEQ Corporation. His recent research efforts have concentrated on computer vision, pattern recognition and image processing of video, aerial photo and LIDAR. He received a Ph.D. in Computer Science and Engineering from the University of Notre Dame. His research at Notre Dame focused on biometrics, including infrared, 2D, and 3D face recognition.