Friday, December 17, 2010

Reading #21: Teddy: A Sketching Interface for 3D Freeform Design (Igarashi)

Summary

The work in this paper a user is asked to draw a 2D sketch which is essentially a 3D projection. Next a 3D model is created from the projection.

Their user study shows that the system is robust and is able to perform the intended task quickly

Discussion

Well it is more a streovision paper than a sketch recognition.

Reading #20: MathPad2: A System for the Creation and Exploration of Mathematical Sketches (LaViola)

Summary

In this paper, Mathpad2 is introduced which lets users to do math. Gestures are used to help segmenting the formulas and to help identification. It is also able to generate graphs and plots. It is also possible to change stroke color to help organization

Discussion

There are definitely room to improve the system such as less reliance on gestures, nevertheless the system looks cool.

Thursday, December 16, 2010

Reading #18: Spatial Recognition and Grouping of Text and Graphics (Shilman)

Summary

In this paper, the authors propose a method of grouping related sketches which does not use timing data.

They essentially search in the space of all possible groupings of strokes, however they incorporate some pruning mechanisms to avoid intractability. One of such pruning is to only group strokes in close proximity. Another measure is restricting the size of such group.

Finally they create an image out of each group and use an adaboost image processing approach to see if it is anything meaningful.

Discussion

This paper is actually taking advantage of the successful application of Viola et al. Adaboost method to face detection. Since it is a very fast algorithm (not in training), they have the freedom of running an expensive search to try different combinations of the sketches to do the grouping.

Reading #17: Distinguishing Text from Graphics in On-line Handwritten Ink (Bishop)

Summary

This system is supposed to distinguish between text and graphics. What they do is to first extract a set of features out of strokes, such as it direction, length width ratio of it, total curvature, etc. Subsequently, it takes uses features of the gaps between strokes and finally by incorporating an HMM, they put all these in context of a sequence.

Their experiments demonstrated the usefulness of temporal context however the advantage of incorporating gap information is not evident in their results

Discussion

As they mention in the paper, they have ignored the length of the gaps, so it would be better to regard stroke within which there is a large gap, independent. However using a temporal model, in this case HMM for the text and graphics seemed reasonable and worked well too

Reading #16: An Efficient Graph-Based Symbol Recognizer (Lee)

Summary

This paper talk about a graph-based symbol recognizer. The authors used a graph called Attribute Relational Graph. The authors try to use such graph and find isomorphism.

The graph represents the topology of the primitive shapes in the graph. In their user study they collected several types of symbols and used their four matching algorithms on the data. The results were from around 68% to 98%

Discussion

I think they could have used better search approaches to this problem

Reading #15: An Image-Based, Trainable Symbol Recognizer for Hand-drawn Sketches (Kara)

Summary

A multi stroke hand drawn symbol recognizer is proposed here. Their technique shows a way to learn symbols definitions using prototype examples allowing users to train new symbols.

Their method has two steps. First, polar coordinates are used to determine angular alignment and eliminate unlikely definitions. Next, the surviving definitions are tested using the normal screen coordinates with four template classifiers. The results of individual classifiers are combined to produce the recognizer’s final decision.

They test their system using two user study. First, on numeric digits where they performed only slightly worse than dedicated recognizers. The second test was in the engineering symbols domain such as resistors, transistors, integral symbol, etc. Their accuracy for being in the top 2 of the result was above 96% in all the tests they conducted in this domain.

Discussion

This paper used some interesting distance measures and ideas which I think I could use earlier in my projects if I had known them.

Reading #14. Using Entropy to Distinguish Shape Versus Text in Hand-Drawn Diagrams (Bhat)

Summary

In this paper entropy rates are proposed as a discriminator between text and shape. Since text strokes usually show more curvature change, the authors have utilized this to distinguish them from other shapes.

Their accuracy of classification reached around 92% which is fine for a relatively simple method.

Discussion

It is a more decent an different approach than the decision tree which was discussed in previous papers.