A Hebrew Tree Bank Based on Cantillation Marks
暂无分享,去创建一个
In the Masoretic text of the Hebrew Bible (HB), the cantillation marks function like a punctuation system that shows the division and subdivision of each verse, forming a tree structure which is similar to the prosodic tree in modern linguistics. However, in the Masoretic text, the structure is hidden in a complicated set of diacritic symbols and the rich information is accessible only to a few trained scholars. In order to make the structural information available to the general public and to automatic processing by the computer, we built a tree bank where the hierarchical structure of each HB verse is explicitly represented in XML format. We coded the punctuation system in a context-tree grammar which was then used by a CYK parser to automatically generate trees for the whole HB. The results show that (1) the CFG correctly encoded the annotation rules and (2) the annotation done by the Masoretes is highly
[1] James D. Price. The syntax of masoretic accents in the Hebrew Bible , 1994 .
[2] Elisabeth Selkirk,et al. Phonology and Syntax: The Relation between Sound and Structure , 1984 .
[3] Joshua R. Jacobson,et al. Chanting the Hebrew Bible , 2005 .