Friday, September 29, 2006

Projector-Guided Painting


Project Description:
This paper presents a novel interactive system for guiding artists to paint using traditional media and tools. The enabling technology is a multi-projector display capable of controlling the appearance of an artist.s canvas. Artists are guided by this display-on-canvas to paint according to a process model we designed to solve 3 common problems with novice painters. The artist paints according to a linear process of painting in layers and, within each layer, a set of colors. Each component of our model of the painting process has an associated interaction mode. Preview mode shows the entire layer as the current painting goal. Blank mode reveals the state of the painting. Color selection mode displays where to paint a target color. Color mixing mode shows how to mix it and orientation mode shows how to paint it. These interaction modes enable the novice to focus on painting sub-tasks in order to simplify the painting process while providing technical guidance ranging from high-level composition to detailed brushwork. We present results of a user study that quantify the benefit that our system can provide to a novice painter.
Publication:
This work will be published and presented at User Interface Software Technology in Montreux, Switzerland in October, 2006.

[Link]

Wednesday, September 27, 2006

[Robotics Institute Thesis Oral 2 Oct 2006]Holistic Modeling and Tracking of Road Scenes

John Sprouse
Robotics Institute
Carnegie Mellon University

Place and Time
NSH 3305
11:00 AM

Abstract
This thesis proposal addresses the problem of road scene understanding for driver warning systems in intelligent vehicles, which require a model of cars, pedestrians, the lane structure of the road, and any static obstacles on it in order to accurately predict possible dangerous situations. Previous work on using computer vision in intelligent vehicle applications stops short of holistic modeling of the entire road scene. In particular, no lane tracking systems exists which detect and track multiple lanes or integrate lane tracking with tracking of cars, pedestrians, and other relevant objects. In this thesis, we focus on the goal of holistic road scene understanding, and we propose contributions in three areas: (1) the low-level detection of road scene elements such as tarmac and painted stripes; (2) modeling and tracking of complex lane structures, and (3) the integration of lane structure tracking with car and pedestrian tracking.

Further Details
A copy of the thesis oral document can be found at
link

Thesis Committee
Takeo Kanade, Chair
Charles Thorpe
Alexei Efros
Simon Baker, Microsoft Research, Seattle

[Robotics Institute Seminar] Is the human hand dexterous because, or in spite, of its anatomical complexity?

Faculty Candidate Talk
Francisco Valero-Cuevas
Cornell University

Time and Place
Mauldin Auditorium (NSH 1305)
Refreshments 3:15 pm
Talk 3:30 pm

Abstract
The human hand is a pinnacle of mechanical versatility unequaled by electromechanical systems. It is clearly a product of brain-body coevolution. However, its anatomical structure shares numerous features with other species. In my work, I explore how the human hand meets the necessary and sufficient mechanical requirement for manipulation. This allows us to begin to distinguish and contrast the complementary contributions of anatomy and the nervous system in order to improve hand rehabilitation and suggest avenues to build better machines.

Speaker Biography
I attended Swarthmore College from 1984-88 where I obtained a BS degree in Engineering. After spending a year in the Indian subcontinent as a Thomas J Watson Fellow, I joined Queen's University in Ontario and worked with Dr. Carolyn Small. The research for my Masters Degree in Mechanical Engineering at Queen's focused on developing non-invasive methods to estimate the kinematic integrity of the wrist joint. In 1991 I joined the doctoral program in the Design Division of the Mechanical Engineering Department at Stanford University. I worked with Dr. Felix Zajac developing a realistic biomechanical model of the human digits. This research, done at the Rehabilitation R & D Center in Palo Alto, focused on predicting optimal coordination patterns of finger musculature during static force production. After completing my doctoral degree in 1997, I joined the core faculty of the Biomechanical Engineering Division at Stanford University as a Research Associate and Lecturer. My research then focused on developing experimental methods to optimize the surgical restoration of hand function following spinal cord injury and peripheral nerve injuries. In 1999 I joined the faculty of the Sibley School of Mechanical and Aerospace Engineering as an Assistant Professor. I also have close ties with the Hospital for Special Surgery in New York City.

Speaker Appointments
For appointments, please contact Jean Harpley(jean@cs.cmu.edu - 8-3802)

[Thesis Proposal] Policies based on Trajectory Libraries

Martin Stolle
Robotics Institute
Carnegie Mellon University

Place and Time
NSH 3305
10:00 AM

Abstract
I present a control approach that uses a library of trajectories to establish a global control law or policy. This is an alternative to methods for finding global policies based on value functions using dynamic programming and also to using plans based on a single desired trajectory. Our method has the advantage of providing reasonable policies much faster than dynamic programming can provide an initial policy. It also has the advantage of providing more robust and global policies than following a single desired trajectory. Trajectory libraries can be created for robots with many more degrees of freedom than what dynamic programming can be applied to as well as for robots with dynamic model discontinuities. Results are shown for the “Labyrinth” marble maze and the Little Dog quadruped robot. The marble maze is a difficult task which requires both fast control as well as planning ahead. In the Little Dog terrain, a quadruped robot has to navigate quickly across small-scale rough terrain. In our past work, I have used global state to represent the knowledge in the trajectory libraries. In order to broaden the use of a library, I propose the use of local state representations, which allow the knowledge represented by a library to be used in novel situations. Three different mechanisms for this transfer are proposed: Information about the goal of a task can be explicitly represented in the local state. Libraries using this representation can be transferred directly to new tasks. Alternatively, the local state representation might not include a goal feature. When using such a library, a search over actions in the library has to be used to pick actions that obtain the goal. Finally, one can cluster the actions in the library in order to create abstract actions. This will simplify the search process.

Further Details
A copy of the thesis proposal document can be found at http://gs3020.sp.cs.cmu.edu/~mstoll/proposal.pdf.

Thesis Committee
Christopher Atkeson, Chair
James Kuffner
Drew Bagnell
Riger Dillmann, University of Karlsruhe

Tuesday, September 26, 2006

Lab Meeting 29 Sep.,2006 (Chihao): 3D Sound Source Localization System Based on Learning of Binaural Hearing

Title: 3D Sound Source Localization System Based on Learning of Binaural Hearing
Author: Hiromichi Nakashima, Toshiharu Mukai
This paper appears in: IEEE SMC 2005 (IEEE International Conference on Systems, Man, and Cybernetics)
Abstract:
We have thus far developed two types of sound source localization system, one of which can localize the horizontal direction and the other the vertical direction. These systems can acquire the localization ability by self-organization through repetition of movement and perception. In this paper, we report a newly built sound source localization system that can detect the direction of a sound source arbitrarily located in front of it. This system is composed of a robot that has two microphones with reflectors corresponding to human’s pinnas. To acquire the horizontal direction, the interaural time difference is used as the auditory cue. To acquire the vertical direction, the features on the audio spectrum induced by the reflectors are used as the auditory cue. The robot can establish the relationship between the cues and the sound direction through learning.

Link

Lab Meeting 29 Sep.,2006 (Vincent): Active Appearance Models

Title : Active Appearance Models

Author : Timothy F. Cootes, Gareth J. Edwards, and Christopher J. Taylor

Origin :
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 23, NO. 6, JUNE 2001

Abstact :
We describe a new method of matching statistical models of appearance to images. A set of model parameters control modes of shape and gray-level variation learned from a training set. We construct an efficient iterative matching algorithm by learning the relationship between perturbations in the model parameters and the induced image errors.

You can find the full article here.

CMU ML talk: Learning-based Deformable Neuroimage Registration

Leonid Teverovskiy, MLD, CMU.
http://www.cs.cmu.edu/~leonid/

September 25

Abstract:
Deformable neuroimage registration is an active and challenging research area. It forms a crucial component of many computational and clinical neuroscience applications, including computer aided diagnosis, statistical quantification of human brain, and atlas-based neuroimage segmentation.

Maximizing the number of correctly estimated voxel correspondences enhances the accuracy of a deformable registration algorithm. Most existing feature-based deformable registration algorithms use a pre-defined set of image features to estimate correspondences for all voxels. These methods have two main weaknesses. First, the featurevector is constructed by the authors of the algorithms rather than automatically selected to minimize registration error. Second, the samefeature vector is used for all the voxels in the whole brain image, without consideration given to the inhomogeneity of the anatomical structures and their corresponding voxels.

We propose a new learning-based deformable registration algorithm that performs feature selection for every voxel. Our algorithm can be trained to accurately register specific anatomical structures as well as the entire neuroimages of specific patient groups. The main novelty of our approach is that it automatically learns feature vectors for distinguished individual image voxels, thus increasing correspondence estimation accuracy. Our method utilizes a decision theoretic approach to systematically calculate the expected correspondence estimation error for a voxel in many different feature spaces, and then select the space with the smallest error. Our feasibility study on the 2D midsagittal slices shows that learning feature subspace increases number ofcorrectly estimated correspondences by 20%.

We will quantitatively evaluate the performance of our deformable registration algorithm and apply it to several medical image analysis problems.

CMU Intelligence Seminar: Semantic Models of Shape

http://www.cs.cmu.edu/~iseminar/

Jovan Popovic', CSAIL, MIT
Tuesday 9/26

Abstract: Conventional representations of shape (splines, meshes, etc.) provide general modeling controls without differentiating between real and meaningless outcomes. This burdens human operators and computational techniques with the task of searching through a vast and cluttered design space. Semantic representations clear up the clutter by attaching human understanding to computational representations of shape and motion.

Bio: Jovan Popovic' is an Associate Professor in the Department of Electrical Engineering and Computer Science, and a member of the Computer Graphics Group in the Computer Science and Artificial Intelligence Laboratory, at the Massachusetts Institute of Technology. Before arriving at MIT in the 2001, Jovan Popovic' received his Ph.D. in Computer Science from Carnegie Mellon University and his B.S. degrees in Mathematics and Computer Science from Oregon State University. His research employs computer science, mathematics and physics to explore the applications of geometric modeling and computer animation to the fields of computer graphics, human-computer interaction, biomechanics, robotics, and computational design.

Monday, September 25, 2006

Call for papers: JFR Special Issue: Safety, Security and Rescue Robots

The Journal of Field Robotics (JFR) announces a special issue on robotic aspects of safety, security, and rescue to examine issues related to the fielding of robots to respond to or prevent emergencies of either natural or man-made origins. Such emergencies provide many challenges for mechanisms, navigation, sensing, networking, collaboration, human/machine interaction, and decision making.

We invite papers that exhibit state-of-the-art theory and methods applied to fielded studies including:

* novel locomotion mechanisms for rough terrain
* robotic sensors and sensing techniques for unstructured or semi-structured terrain
* novel human/robot interaction devices and paradigms for emergency response
* collaborative systems of land/sea/air vehicles for search/rescue/assessment
* lessons learned from robotic land/sea/air deployments in mines, collapsed structures, and wide area disasters

The complete call for papers for this special issue can be found at:
http://journalfieldrobotics.org/si.html

Please note that the deadline for submissions is November 1, 2006.

We look forward to your submission.

-Richard Voyles and Howie Choset

Sunday, September 24, 2006

News: Robot manufacturing to be Taiwan's next booming industry: MOEA

The China Post
2006/9/23
TAIPEI, CNA

The research and development of artificial intelligence robots could help the total output value of Taiwan's machinery industry to top NT$1 trillion (US$30.4 billion) by 2009, with an annual growth rate of no less than 30 percent before 2012, the Industrial Development Bureau (IDB) under the Ministry of Economic Affairs said yesterday.

The IDB also predicted that by 2016, A.I. robot manufacture alone could generate an output value of NT$250 billion and an export value of NT$175 billion, as well as contribute at least 1.35 percent to Taiwan's GDP, adding that 22,000 jobs could be created.

Moreover, the IDB estimated that the global production value of the robot industry could exceed that of the automobile sector worldwide by 2020, reaching some US$1.4 trillion.

...... See the full article.

Saturday, September 23, 2006

Lab Meeting 29 Sep.,2006 (Ashin):Using GPS to learn significant locations and predict movement across multiple users

Authors:Daniel Ashbrook and Thad Starner

From:Personal and Ubiquitous Computing Volume 7, Number 5 / October, 2003

Abstract:
Wearable computers have the potential to act as intelligent agents in everyday life and assist the user in a variety of tasks, using context to determine howto act. Location is the most common form of context used by these agents to determine the user's task.However, another potential use of location context isthe creation of a predictive model of the user's future movements. We present a system that automatically clusters GPS data taken over an extended periodof time into meaningful locations at multiple scales.These locations are then incorporated into a Markovmodel that can be consulted for use with a variety ofapplications in both single user and collaborative scenarios.

[Link]

Thursday, September 21, 2006

IEEE news: How to manage an impending deluge of new data?

The WalMart project, which aims to have every item delivered and sold tagged with a radio frequency identification strip (RFID), is only the tip of an information iceberg. As the data streaming from computers, sensors, and other real-time devices swells, a new infrastructure will be required to manage it. Software engineers are developing stream-processing engines to handle the flood.

See "Data Torrents and Rivers," by Michael Stonebraker: the link

IEEE news: Sounding out IEEE's fellows

How do some of the IEEE's most distinguished and accomplished members see the future of technology? Spectrum polled 700 of the organization's fellows to find out. A third think we'll eventually have 3D televisions in our homes, but more than two fifths doubt anybody will ever market a quantum computer. Majorities believe that microscale robots will be a reality and that computers will soon be able to process speech and writing with almost perfect accuracy. But there's deep skepticism about "cold fusion" and room-temperature superconductors.

See "Bursting Tech Bubbles Before They Balloon," by Marina Gorbis and David Pescovitz: the link.

Wednesday, September 20, 2006

Lab meeitng 22 Sep., 2006(ZhenYu):Communication Robots for Elementary Schools

Authors: Takayuki Kanda,Hiroshi Ishiguro

From: Proc. AISB'05 Symposium Robot Companions: Hard Problems and Open Challenges in Robot-Human Interaction, pp. 54-63, April 2005.

Abstract: This paper reports our approaches and efforts for developing communication robots for elementary schools. In particular, we describe the fundamental mechanism of the interactive humanoid robot, Robovie, for interacting with multiple persons, maintaining relationships, and estimating social relationships among children. The developed robot Robovie was applied for two field experiments at elementary schools. The first experiment purpose using it as a peer tutor of foreign language education, and the second was purposed for establishing longitudinal relationships with children. We believe that these results demonstrate a positive perspective for the future possibility of realizing a communication robot that works in elementary schools.


[Link]

Tuesday, September 19, 2006

Lab meeitng 22 Sep., 2006 (Eric): Dual Photography

Authors: Pradeep Sen, Billy Chen, Gaurav Garg, Stephen R. Marschner, Mark Horowitz, Marc Levoy, Hendrik P. A. Lensch

From: ACM SIGGRAPH 2005 conference proceedings

Abstract: We present a novel photographic technique called dual photography,
which exploits Helmholtz reciprocity to interchange the lights
and cameras in a scene. With a video projector providing structured
illumination, reciprocity permits us to generate pictures from
the viewpoint of the projector, even though no camera was present
at that location. The technique is completely image-based, requiring
no knowledge of scene geometry or surface properties, and
by its nature automatically includes all transport paths, including
shadows, inter-refections and caustics. In its simplest form, the
technique can be used to take photographs without a camera; we
demonstrate this by capturing a photograph using a projector and
a photo-resistor. If the photo-resistor is replaced by a camera, we
can produce a 4D dataset that allows for relighting with 2D incident
illumination. Using an array of cameras we can produce a 6D
slice of the 8D re-refectance feld that allows for relighting with arbitrary
light felds. Since an array of cameras can operate in parallel
without interference, whereas an array of light sources cannot, dual
photography is fundamentally a more effcient way to capture such
a 6D dataset than a system based on multiple projectors and one
camera. As an example, we show how dual photography can be
used to capture and relight scenes.

[Link]

Lab meeitng 22 Sep., 2006 (Bright): A Novel System for Tracking Pedestrains Using Multiple Single-Row Laser-Range Scanners

Authors: Huijing Zhao and Ryosuke Shibasaki

From: IEEE TRANSACTIONS ON SYSTEMS, MAN, AND CYBERNETICS—PART A: SYSTEMS AND HUMANS, VOL. 35, NO. 2, MARCH 2005

Abstract: In this research, we propose a novel system for
tracking pedestrians in a wide and open area, such as a shopping
mall and exhibition hall, using a number of single-row laser-range
scanners (LD-A), which have a profiling rate of 10 Hz and a
scanning angle of 270 . LD-As are set directly on the floor doing
horizontal scanning at an elevation of about 20 cm above the
ground, so that horizontal cross sections of the surroundings,
containing moving feet of pedestrians as well as still objects, are
obtained in a rectangular coordinate system of real dimension.
The data of moving feet are extracted through background subtraction
by the client computers that control each LD-A, and sent
to a server computer, where they are spatially and temporally integrated
into a global coordinate system. A simplified pedestrian’s
walking model based on the typical appearance of moving feet is
defined and a tracking method utilizing Kalman filter is developed
to track pedestrian’s trajectories. The system is evaluated through
both real experiment and computer simulation. A real experiment
is conducted in an exhibition hall, where three LD-As are used
covering an area of about 60 x 60 m2. Changes in visitors’ flow
during the whole exhibition day are analyzed, where in the peak
hour, about 100 trajectories are extracted simultaneously. On the
other hand, a computer simulation is conducted to quantitatively
examine system performance with respect to different crowd
density.

Link

CMU thesis proposal: Scale Selection and Invariance in Low Level Vision

Ranjith Unnikrishnan, Robotics Institute
18 Sep 2006

The representation of objects through locally computed features is a concept common to many approaches in 2-D and 3-D computer vision. The use of local information to infer global properties aims to serve several purposes such as robustness to outlying structures, variation in viewing conditions, noise and occlusion. Reliable computation of relevant local attributes is thus an important part of any practical vision system intended to perform higher level reasoning.

The task of making such local observations necessitates making choices of the neighborhood size within which the computation is performed, also referred to as the /scale/ of the observation. This in turn poses several unanswered questions relevant to both data representation (e.g. reconstruction and compression) as well as data identification (e.g. object detection and classification). At what scale is it meaningful to compute a local feature? What is the optimal neighborhood size for estimating local geometric properties from data? While many advances have been made in a theory of scale for 2-D luminance images, little attention has been paid to the domains of unorganized point clouds (as would be acquired with a laser range scanner) or to alternate representations of images (such as color or other pixel-wise functions such as optical flow).

This thesis explores the problems of scale selection and invariance in previously unaddressed problem domains, and proposes solutions for several useful vision tasks:

* We propose to extend current application of scale theory for interest region extraction in 2-D images to alternate, potentially more useful representations. As an example, we demonstrate how both scale as well as illuminant invariant keypoint detection may be achieved in the case of color (RGB) images without having to estimate the properties of the illuminant.

* We present methods to robustly compute local differential properties from non-uniform unstructured point clouds. In particular, we show how data-driven adaptation of the neighborhood size in local PCA when computing tangents (normals) from spatial curves (surfaces) can make even this naive estimator more robust than leading fixed-scale alternatives.

* We propose the development of new scale-space representations of 3-D point cloud data that are robust to changes in sampling. By this, we advocate changing the current practice of using a single globally fixed value of scale when computing shape descriptors from 3-D data to that of using a value that is locally data-driven.

* We propose to investigate the application of local intrinsic scale detection for manifold learning. The aims of this analysis are improved statistical properties and robustness of embedding functions and regularizers to sampling variations in the dataset.

Overall, the expected contributions of this thesis are new technique and tools for the scale selection problem that is fundamental to local data analysis and learning from real-world measurements.

Further Details: A copy of the thesis proposal document can be found at the link.

CMU vasc talk: Visual Recognition and Tracking for Perceptive Interfaces

Trevor Darrell
MIT CSAIL
http://people.csail.mit.edu/trevor/

Devices should be perceptive, and respond directly to their human user and/or environment. In this talk I'll present new computer vision algorithms for fast recognition, indexing, and tracking that make this possible, enabling multimodal interfaces which respond to users' conversational gesture and body language, robots which recognize common object categories, and mobile devices which can search using visual cues of specific objects of interest. As time permits, I'll describe recent advances in real-time human pose tracking for multimodal interfaces, including new methods which exploit fast computation of approximate likelihood with a pose-sensitive image embedding. I'll also present our linear-time approximate correspondence kernel, the Pyramid Match, and its use for image indexing and object recognition, and discovery of object categories. Throughout the talk, I'll show interface examples including grounded multimodal conversation as well as mobile image-based information retrieval applications based on these techniques.

BIO: Trevor Darrell is an Associate Professor of Electrical Engineering and Computer Science at M.I.T. He leads the Vision Interface Group at the Computer Science and Artificial Intelligence Laboratory. His interests include computer vision, interactive graphics, and machine learning. Prior to joining the faculty of MIT he worked as a Member of the Research Staff at Interval Research in Palo Alto, CA, researching vision-based interface algorithms for consumer applications. He received his PhD and SM from MIT in 1996 and 1991, respectively, while working at the Media Laboratory, and the BSE from the University of Pennsylvania in 1988, where he worked in the GRASP Robotics Laboratory.

CMU ML talks: UAI 2006 conference review

1. On the Number of Samples Needed to Learn the Correct Structure of a Bayesian Network

by Or Zuk, Shiri Margel and Eytan Domany

Bayesian Networks (BNs) are useful tools giving a natural and compact representation of joint probability distributions. In many applications one needs to learn a Bayesian Network (BN) from data. In this context, it is important to understand the number of samples needed in order to guarantee a successful learning. Previous works have studied BNs sample complexity, yet they mainly focused on the requirement that the learned distribution will be close to the original distribution which generated the data. In this work, we study a different aspect of the learning task, namely the number of samples needed in order to learn the correct structure of the network. We give both asymptotic results (lower and upper-bounds) on the probability of learning a wrong structure, valid in the large sample limit, and experimental results, demonstrating the learning behavior for feasible sample sizes.


2. Non-Minimal Triangulations for Mixed Stochastic/Deterministic Graphical Models
by Chris D. Bartels and Jeff A. Bilmes

We observe that certain large-clique graph triangulations can be useful for reducing computational requirements when making queries on mixed stochastic/deterministic graphical models. We demonstrate that many of these large clique triangulations are non-minimal and are thus unattainable via the elimination algorithm. We introduce ancestral pairs as the basis for novel triangulation heuristics and prove that no more than the addition of edges between ancestral pairs need be considered when searching for state space optimal triangulations in such graphs. Empirical results on random and real world graphs are given. We also present an algorithm and correctness proof for determining if a triangulation can be obtained via elimination, and we show that the decision problem associated with finding optimal state space triangulations in this mixed setting is NP-complete.

Monday, September 18, 2006

Lab meeting this Fall

Bright, Eric and Zhen-Yu are the speakers of this week. Please post your talks asap.

When: 10:30 AM - 12:30 PM
Where: CSIE R424/426

Best,

-Bob