Among the various proposals answering the shortcomings of Document Type Definitions (DTDs), XML Schema is the most widely used. Although DTDs and XML Schema Definitions (XSDs) differ syntactically, they are still quite related on an abstract level. Indeed, freed from all syntactic sugar, XML Schemas can be seen as an extension of DTDs with a restricted form of specialization. In the present paper, we inspect a number of DTDs and XSDs harvested from the web and try to answer the following questions: (1) which of the extra features/expressiveness of XML Schema not allowed by DTDs are effectively used in practice; and, (2) how sophisticated are the structural properties (i.e. the nature of regular expressions) of the two formalisms. It turns out that at present real-world XSDs only sparingly use the new features introduced by XML Schema: on a structural level the vast majority of them can already be defined by DTDs. Further, we introduce a class of simple regular expressions and obtain that a surprisingly high fraction of the content models belong to this class. The latter result sheds light on the justification of simplifying assumptions that sometimes have to be made in XML research.
[1]
Arvind Malhotra,et al.
Xml schema part 2: datatypes
,
1999
.
[2]
David C. Fallside,et al.
Xml schema part 0: primer
,
2000
.
[3]
Thomas Schwentick,et al.
Complexity of Decision Problems for Simple Regular Expressions
,
2004,
MFCS.
[4]
C. M. Sperberg-McQueen,et al.
Extensible Markup Language (XML)
,
1997,
World Wide Web J..
[5]
Yannis Papakonstantinou,et al.
DTD inference for views of XML data
,
2000,
PODS.
[6]
Arnaud Sahuguet.
Everything You Ever Wanted to Know About DTDs, But Were Afraid to Ask
,
2000,
WebDB.
[7]
Byron Choi,et al.
What are real DTDs like?
,
2002,
WebDB.
[8]
Murali Mani,et al.
Taxonomy of XML schema languages using formal language theory
,
2005,
TOIT.
[9]
B. E. Eckbo,et al.
Appendix
,
1826,
Epilepsy Research.
[10]
Derick Wood,et al.
Regular Tree Languages Over Non-Ranked Alphabets
,
1998
.
[11]
Nils Klarlund,et al.
Document Structure Description 1.0
,
2000
.
[12]
Derick Wood,et al.
One-Unambiguous Regular Languages
,
1998,
Inf. Comput..
[13]
Eric van der Vlist,et al.
XML Schema
,
2002
.