Main Article Content
Abstract
This paper resolves the trouble needed in detecting Gurmukhi and Hindi text object present in images. As there are lots of differences among the features of script in English and Indian languages (for example line over text string), we cannot directly apply the existing algorithms. Therefore in this paper we propose a new region based bottom up method to extract text from images, where we first identify elementary substructures using connected component and edges, and then merge them successively into huge buildings until all book areas are recognized. We localized text candidates by extracting closed boundaries. As each character counter has high contrast as compared to its neighbors, all character pixels and several non-character pixels which exhibit excessive neighborhood intensity difference are localized as written text within the advantage image.