Abstract.
Finger spelling is an art of communicating by signsmade with fingers, and has been introduced into sign language to serveas a bridge between the sign language and the verbal language.Previous approaches to finger spelling recognition are classified intotwo categories: glove-based and vision-based approaches. The glove-based approach is simpler and more accurate recognizing work of hand posture than vision-based, yet the interfaces require the user to wear a cumbersome and carry a load of cables that connected the device to a computer. In contrast, the vision-based approaches provide an attractive alternative to the cumbersome interface, and promise more natural and unobtrusive human-computer interaction. The vision-based approaches generally consist of two steps: hand extraction and recognition, and two steps are processed independently. This paper proposes real-time vision-based Korean finger spelling recognition system by integrating hand extraction into recognition. First, we tentatively detect a hand region using CAMShift algorithm.Then fill factor and aspect ratio estimated by width and height of detected hand regions are used to choose candidate from database,which can reduce the number of matching in recognition step. Torecognize the finger spelling, we use DTW(dynamic time warping) based on modified chain codes, to be robust to scale and orientation variations. In this procedure, since accurate hand regions, without holes and noises, should be extracted to improve the precision, we use graph cuts algorithm that globally minimize the energy function elegantly expressed by Markov random fields (MRFs). In the experiments, the computational times are less than 130 ms, and the times are not related to the number of templates of finger spellings in database, as candidate templates are selected in extraction step.
Keywords: CAMShift, DTW, Graph Cuts, and MRF.
click
http://hci.ssu.ac.kr/ajpark/[IC]GestureRecognition.pdf
to download the paper.
2008.
2008-10-01
[IC] Graph Cuts-based Automatic Color Image Segmentation using Mean Shift Analysis
Abstract.
A graph cuts method has recently attracted a lot of attention for image segmentation, as it can minimize an energy function composed of data term estimated in feature space and smoothness term estimated in an image domain. Although previous approaches using graph cuts have shown good performance for image segmentation, they manually obtained prior information to estimate the data term, thus automatic image segmentation is one of issues in application using the graph cuts method. To automatically estimate the data term, GMM (Gaussian mixture model) is generally used, but it is practicable only for classes with a hyper-spherical or hyper-ellipsoidal shape, as the class was represented based on the covariance matrix centered on the mean. For arbitrary-shaped classes, this paper proposes graph cuts-based image segmentation using mean shift analysis. As prior information to estimate the data term, we use the set of mean trajectories toward each mode from initial means randomly selected in L*u*v* feature space. Since the mean shift procedure requires many computational times, we transform features in continuous feature space into 3D discrete grid, and use 3D kernel based on the first moment in the grid, which are needed to move the means to modes. In the experiments, we investigated problems of normalized cuts-based and mean shift-based segmentation and graph cuts-based segmentation using GMM. As a result, the proposed method showed better performance than previous three methods on Berkeley segmentation dataset.
click
http://hci.ssu.ac.kr/ajpark/[IC]ColorImageSegmentation.pdf
to download the paper.
2008.
A graph cuts method has recently attracted a lot of attention for image segmentation, as it can minimize an energy function composed of data term estimated in feature space and smoothness term estimated in an image domain. Although previous approaches using graph cuts have shown good performance for image segmentation, they manually obtained prior information to estimate the data term, thus automatic image segmentation is one of issues in application using the graph cuts method. To automatically estimate the data term, GMM (Gaussian mixture model) is generally used, but it is practicable only for classes with a hyper-spherical or hyper-ellipsoidal shape, as the class was represented based on the covariance matrix centered on the mean. For arbitrary-shaped classes, this paper proposes graph cuts-based image segmentation using mean shift analysis. As prior information to estimate the data term, we use the set of mean trajectories toward each mode from initial means randomly selected in L*u*v* feature space. Since the mean shift procedure requires many computational times, we transform features in continuous feature space into 3D discrete grid, and use 3D kernel based on the first moment in the grid, which are needed to move the means to modes. In the experiments, we investigated problems of normalized cuts-based and mean shift-based segmentation and graph cuts-based segmentation using GMM. As a result, the proposed method showed better performance than previous three methods on Berkeley segmentation dataset.
click
http://hci.ssu.ac.kr/ajpark/[IC]ColorImageSegmentation.pdf
to download the paper.
2008.
[IC] Neural Network Implementation using CUDA and OpenMP
Abstract.
Many algorithms for image processing and patternrecognition have recently been implemented on GPU(graphic processing unit) for faster computationaltimes. However, the implementation using GPUencounters two problems. First, the programmershould master the fundamentals of the graphicsshading languages that require the prior knowledge oncomputer graphics. Second, in a job which needs muchcooperation between CPU and GPU, which is usual inimage processings and pattern recognitions contraryto the graphics area, CPU should generate raw featuredata for GPU processing as much as possible toeffectively utilize GPU performance. This paperproposes more quick and efficient implementation ofneural networks on both GPU and multi-core CPU.We use CUDA (compute unified device architecture)that can be easily programmed due to its simple Clanguage-like style instead of GPGPU to solve the firstproblem. Moreover, OpenMP (Open Multi-Processing)is used to concurrently process multiple data withsingle instruction on multi-core CPU, which results ineffectively utilizing the memories of GPU. In theexperiments, we implemented neural networks-basedtext detection system using the proposed architecture,and the computational times showed about 15 timesfaster than implementation using CPU and about 4times faster than implementation on only GPU withoutOpenMP.
click
http://hci.ssu.ac.kr/ajpark/[IC]CUDAforNN.pdf
to download the papaer.
2008.
Many algorithms for image processing and patternrecognition have recently been implemented on GPU(graphic processing unit) for faster computationaltimes. However, the implementation using GPUencounters two problems. First, the programmershould master the fundamentals of the graphicsshading languages that require the prior knowledge oncomputer graphics. Second, in a job which needs muchcooperation between CPU and GPU, which is usual inimage processings and pattern recognitions contraryto the graphics area, CPU should generate raw featuredata for GPU processing as much as possible toeffectively utilize GPU performance. This paperproposes more quick and efficient implementation ofneural networks on both GPU and multi-core CPU.We use CUDA (compute unified device architecture)that can be easily programmed due to its simple Clanguage-like style instead of GPGPU to solve the firstproblem. Moreover, OpenMP (Open Multi-Processing)is used to concurrently process multiple data withsingle instruction on multi-core CPU, which results ineffectively utilizing the memories of GPU. In theexperiments, we implemented neural networks-basedtext detection system using the proposed architecture,and the computational times showed about 15 timesfaster than implementation using CPU and about 4times faster than implementation on only GPU withoutOpenMP.
click
http://hci.ssu.ac.kr/ajpark/[IC]CUDAforNN.pdf
to download the papaer.
2008.
2008-09-30
[IC] Clustering of Trained Self-Organizing Feature Maps based on s-t Graph Cuts
Abstract.
The Self-organizing Feature Map(SOFM) that is one of unsupervised neural networks is a very powerful tool for data clustering and visualization in high-dimensional data sets. Although the SOFM has been applied in many engineering problems, it needs to cluster similar weights into one class on the trained SOFM as a post-processing, which is manually performed in many cases. The traditional clustering algorithms, such as k-means, on the trained SOFM, but do not yield satisfactory results, especially when clusters have arbitrary shapes. This paper proposes automatic clustering on trained SOFM via graph cuts, which can both deal with arbitrary cluster shapes and be globally optimized by graph cuts. When using graph cuts, the graph must have two additional nodes, called terminals, and weights between the terminals and nodes of the graph are generally setting based on data manually obtained by users. The proposed method automatically sets the weights based on mode-seeking on a distance matrix. Experimental results demonstrated the effectiveness of the proposed method in texture segmentation. In the experimental results, the proposed method improved precision rates compared with previous traditional clustering algorithm, as the method can deal with arbitrary cluster shapes based on the graph-theoretic clustering and globally optimize the clustering of the trained SOFM by graph cuts.
click
http://hci.ssu.ac.kr/ajpark/[MLDM]Clustering.pdf
to download the paper.
2007.
The Self-organizing Feature Map(SOFM) that is one of unsupervised neural networks is a very powerful tool for data clustering and visualization in high-dimensional data sets. Although the SOFM has been applied in many engineering problems, it needs to cluster similar weights into one class on the trained SOFM as a post-processing, which is manually performed in many cases. The traditional clustering algorithms, such as k-means, on the trained SOFM, but do not yield satisfactory results, especially when clusters have arbitrary shapes. This paper proposes automatic clustering on trained SOFM via graph cuts, which can both deal with arbitrary cluster shapes and be globally optimized by graph cuts. When using graph cuts, the graph must have two additional nodes, called terminals, and weights between the terminals and nodes of the graph are generally setting based on data manually obtained by users. The proposed method automatically sets the weights based on mode-seeking on a distance matrix. Experimental results demonstrated the effectiveness of the proposed method in texture segmentation. In the experimental results, the proposed method improved precision rates compared with previous traditional clustering algorithm, as the method can deal with arbitrary cluster shapes based on the graph-theoretic clustering and globally optimize the clustering of the trained SOFM by graph cuts.
click
http://hci.ssu.ac.kr/ajpark/[MLDM]Clustering.pdf
to download the paper.
2007.
[IC] Edge-based Eye Region Detection in Rotated Face using Global Orientation Histogram
Abstract.
Automatic human face analysis and recognition have become one of the most important research topics in robot society, and research on automatic eye region detection has recently attracted a lot of attentions, as the most important feature for human faces is eyes. Although much effect has been spent, the problem of automatic eye region detection is still challenging, as most of the existing methods mainly focus on eye detection in the frontal face without consideration of the factors, such as non-frontal faces and lighting conditions. This paper proposes an eye region detection method in faces rotated around the front-to-back axis called rotated faces, based on edge information that shows fast computational times. The proposed method consists of two steps: making frontal faces, and then detecting eye regions. The rotated angle of face regions can be estimated by analyzing histogram accumulated from edge orientation of the face region, called global orientation histogram. The rotated face can be frontal faces based on the estimated angle, and then the eye regions are detected in frontal face by analyzing edge orientation histogram of components grouping adjacent edges, called local orientation histogram, already verified by experiment of previous works. Experiment results demonstrated the effectiveness of the proposed method using 300 face images provided from THE Weizmann Institute of Science, and achieved precision rates of 83.5% and computational times of 0.5 seconds.
click
http://hci.ssu.ac.kr/ajpark/[ICRA]Edgebased.pdf
to download the paper.
2007.
Automatic human face analysis and recognition have become one of the most important research topics in robot society, and research on automatic eye region detection has recently attracted a lot of attentions, as the most important feature for human faces is eyes. Although much effect has been spent, the problem of automatic eye region detection is still challenging, as most of the existing methods mainly focus on eye detection in the frontal face without consideration of the factors, such as non-frontal faces and lighting conditions. This paper proposes an eye region detection method in faces rotated around the front-to-back axis called rotated faces, based on edge information that shows fast computational times. The proposed method consists of two steps: making frontal faces, and then detecting eye regions. The rotated angle of face regions can be estimated by analyzing histogram accumulated from edge orientation of the face region, called global orientation histogram. The rotated face can be frontal faces based on the estimated angle, and then the eye regions are detected in frontal face by analyzing edge orientation histogram of components grouping adjacent edges, called local orientation histogram, already verified by experiment of previous works. Experiment results demonstrated the effectiveness of the proposed method using 300 face images provided from THE Weizmann Institute of Science, and achieved precision rates of 83.5% and computational times of 0.5 seconds.
click
http://hci.ssu.ac.kr/ajpark/[ICRA]Edgebased.pdf
to download the paper.
2007.
Subscribe to:
Posts (Atom)