research-article

Morphology-Based Segmentation Combination for Arabic Mention Detection

Authors:
Yassine Benajiba

Center for Computational Learning Systems, Columbia University

Center for Computational Learning Systems, Columbia University
View Profile

,
Imed Zitouni

IBM T. J. Watson Research Center

IBM T. J. Watson Research Center
View Profile

ACM Transactions on Asian Language Information Processing Volume 8 Issue 4Article No.: 16pp 1–18https://doi.org/10.1145/1644879.1644883

Published:01 December 2009Publication History

ACM Transactions on Asian Language Information Processing

Abstract

The Arabic language has a very rich/complex morphology. Each Arabic word is composed of zero or more prefixes, one stem and zero or more suffixes. Consequently, the Arabic data is sparse compared to other languages such as English, and it is necessary to conduct word segmentation before any natural language processing task. Therefore, the word-segmentation step is worth a deeper study since it is a preprocessing step which shall have a significant impact on all the steps coming afterward. In this article, we present an Arabic mention detection system that has very competitive results in the recent Automatic Content Extraction (ACE) evaluation campaign. We investigate the impact of different segmentation schemes on Arabic mention detection systems and we show how these systems may benefit from more than one segmentation scheme. We report the performance of several mention detection models using different kinds of possible and known segmentation schemes for Arabic text: punctuation separation, Arabic Treebank, and morphological and character-level segmentations. We show that the combination of competitive segmentation styles leads to a better performance. Results indicate a statistically significant improvement when Arabic Treebank and morphological segmentations are combined.

References

Benajiba, Y., Diab, M., and Rosso, P. 2008. Arabic named entity recognition using optimized feature sets. In Proceedings of the Joint Meeting of the Conference on Empirical Methods in Natural Language Processing (EMNLP’08). Google ScholarDigital Library
Berger, A., Della Pietra, S., and Della Pietra, V. 1996. A maximum entropy approach to natural language processing. Comput. Linguist. 22, 1, 39--71. Google ScholarDigital Library
Chen, S. and Rosenfeld, R. 2000. A survey of smoothing techniques for ME models. IEEE Trans. Speech Audio Proc.Google ScholarCross Ref
Chen, S. F. and Goodman, J. 1998. An empirical study of smoothing techniques for language modeling. Tech. rep., TR-10-98, Center for Research in Computing Technology, Harvard University.Google Scholar
Diab, M., Hacioglu, K., and Jurafsky, D. 2004. Automatic tagging of Arabic text: From raw text to base phrase chunks. In Proceedings of the Human Language Technology Conference/North American Chapter of the Association for Computational Linguistics (HLT/NAACL’04). Google ScholarDigital Library
Florian, R., Hassan, H., Ittycheriah, A., Jing, H., Kambhatla, N., Luo, X., Nicolov, N., and Roukos, S. 2004. A statistical model for multilingual entity detection and tracking. In Proceedings of the Human Language Technology Conference/North American Chapter of the Association for Computational Linguistics (HLT/NAACL’04).Google Scholar
Goodman, J. 2002. Sequential conditional generalized iterative scaling. In Proceedings of the Association of Computer Linguistics (ACL’02). Google ScholarDigital Library
Graff, D. 2003. Arabic gigaword. http://www.ldc.upenn.edu/.Google Scholar
Habash, N. and Sadat, F. 2006. Combination of Arabic preprocessing schemes for statistical machine translation. In Proceedings of the Association of Computer Linguistics (ACL’06). Google ScholarDigital Library
Jing, H., Florian, R., Luo, X., Zhang, T., and Ittycheriah, A. 2003. How to get a Chinese Name (Entity): Segmentation and combination issues. In Proceedings of the Joint Meeting of the Conference on Empirical Methods in Natural Language Processing (EMNLP’03). Google ScholarDigital Library
Lafferty, J., McCallum, A., and Pereira, F. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings of the International Conference on Machine Learning (ICML’01). Google ScholarDigital Library
Lee, Y.-S., Papineni, K., Roukos, S., Emam, O., and Hassan, H. 2003. Language model based Arabic word segmentation. In Proceedings of the Association of Computer Linguistics (ACL’03). 399--406. Google ScholarDigital Library
Maamouri, M., Bies, A., Buckwalter, T., and Mekki, W. 2004. The Penn Arabic Treebank: Building a large-scale annotated Arabic corpus. In Proceedings of the NEMLAR Conference on Arabic Language Resources and Tools (NEMLAR’04).Google Scholar
Mohri, M., Pereira, F. C. N., and Riley, M. 1998. A rational design for a weighted finite-state transducer library. In Proceedings of the 2nd International Workshop on Implementation and Application of Automata (CIAA’98). D. Wood and S. Yu, Eds. Lecture Notes in Computer Science, vol. 1436. Springer-Verlag: Berlin. 144--158. Google ScholarDigital Library
NIST. 2007. The ACE evaluation plan. www.nist.gov/speech/tests/ace/index.htm.Google Scholar
Noreen, E. W. 1989. Computer-Intensive Methods for Testing Hypotheses. John Wiley & Sons.Google Scholar
Ramshaw, L. and Marcus, M. 1994. Exploring the statistical derivation of transformational rule sequences for part-of-speech tagging. In Proceedings of the ACL Workshop on Combining Symbolic and Statistical Approaches to Language (ACL’94). 128--135.Google Scholar
Ramshaw, L. and Marcus, M. 1995. Text chunking using transformation-based learning. In Proceedings of the 3rd Workshop on Very Large Corpora (WVLC’95). D. Yarowsky and K. Church, Eds. Association for Computational Linguistics, 82--94.Google Scholar
Tjong Kim Sang, E. F. 2002. Introduction to the CONLL-2002 shared task: Language-independent named entity recognition. In Proceedings of the Conference on Natural Language Learning (CoNLL’02). 155--158. Google ScholarDigital Library
Tjong Kim Sang, E. F. and Veenstra, J. 1999. Representing text chunks. In Proceedings of the Conference on the European Chapter of the Association for Computational Linguistics (EACL’99). Google ScholarDigital Library
Toutanova, K., Klein, D., Manning, C., and Singer, Y. 2003. Feature-rich part-of-speech tagging with a cyclic dependency network. In Proceedings of the Human Language Technology Conference/North American Chapter of the Association for Computational Linguistics (HLT/NAACL’03). Google ScholarDigital Library
Zitouni, I., Sorensen, J., Luo, X., and Florian, R. 2005. The impact of morphological stemming on Arabic mention detection and conference resolution. In Proceedings of the ACL Workshop on Computing Approaches to Semitic Languages (CASL’05). Google ScholarDigital Library

Recommendations

Cross-Language Information Propagation for Arabic Mention Detection

In the last two decades, significant effort has been put into annotating linguistic resources in several languages. Despite this valiant effort, there are still many languages left that have only small amounts of such resources. The goal of this article ...
Read More
Segmentation of Arabic Handwriting Based on both Contour and Skeleton Segmentation
ICDAR '09: Proceedings of the 2009 10th International Conference on Document Analysis and Recognition

We propose a new algorithm for segmentation of off-line handwritten Arabic words. The algorithm segments the connected letters to smaller segments each of which contains no more than three letters. Each letter may be segmented to at most five pieces. In ...
Read More
A survey on Arabic character segmentation

Arabic character segmentation is a necessary step in Arabic Optical Character Recognition (OCR). The cursive nature of Arabic script poses challenging problems in Arabic character recognition; however, incorrectly segmented characters will cause ...
Read More

Comments

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Article

Published in

ACM Transactions on Asian Language Information Processing Volume 8, Issue 4
December 2009
121 pages
ISSN:1530-0226
EISSN:1558-3430
DOI:10.1145/1644879
Issue’s Table of Contents

Copyright © 2009 ACM
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]
Sponsors
In-Cooperation
Publisher
Association for Computing Machinery
New York, NY, United States
Publication History
- Published: 1 December 2009
- Accepted: 1 September 2009
- Revised: 1 August 2009
- Received: 1 March 2009
Published in talip Volume 8, Issue 4

Permissions
Request permissions about this article.
Request Permissions

Check for updates
Author Tags
Arabic information extraction
Arabic mention detection
Arabic segmentation
Qualifiers
- research-article
- Research
- Refereed
Conference
Funding Sources
Other Metrics
View Article Metrics

Article Metrics
- 4
  Total Citations
  View Citations
- 326
  Total Downloads
- Downloads (Last 12 months)2
- Downloads (Last 6 weeks)0
Other Metrics
View Author Metrics
Cited By
View all

PDF Format

View or Download as a PDF file.

PDF

eReader

View online with eReader.

eReader

Morphology-Based Segmentation Combination for Arabic Mention Detection

ACM Transactions on Asian Language Information Processing

Abstract

References

Cited By

Recommendations

Cross-Language Information Propagation for Arabic Mention Detection

Segmentation of Arabic Handwriting Based on both Contour and Skeleton Segmentation

A survey on Arabic character segmentation

Comments

Login options

Full Access

Published in

Sponsors

In-Cooperation

Publisher

Publication History

Permissions

Check for updates

Author Tags

Qualifiers

Conference

Funding Sources

Other Metrics

Article Metrics

Other Metrics

Cited By

PDF Format

eReader

Digital Edition

Caption

Morphology-Based Segmentation Combination for Arabic Mention Detection

ACM Transactions on Asian Language Information Processing

Abstract

References

Cited By

Recommendations

Cross-Language Information Propagation for Arabic Mention Detection

Segmentation of Arabic Handwriting Based on both Contour and Skeleton Segmentation

A survey on Arabic character segmentation

Comments

Login options

Full Access

Published in

Sponsors

In-Cooperation

Publisher

Publication History

Permissions

Check for updates

Author Tags

Qualifiers

Conference

Funding Sources

Article Metrics

Other Metrics

PDF Format

eReader

Digital Edition

Share this Publication link

Share on Social Media