HSEIWP 96.24

``Theory-Neutral'' Syntactic Tagging of Text

Hebrew Syntax Encoding Initiative Working Papers, no. 96.2
(revised fourth draft, August 1997)

Vincent DeCaen 
University of Toronto

* * *

Contents

1.   Preamble

Part I: The Logic of Syntax
2.   Introduction
3.   Syntactic Metatheory
4.   Bases of Syntactic Explanation

Part II: Three Bases of Syntactic Explanation
5.   The Skeleton of Phrase Structure
6.   Case Annotation
7.   Thematic Roles

Part III: First Approximation
8. Sample Parsing
9. Remaining Questions and Conclusion

Bibliography

* * *

1. Preamble

1.1  This working paper, HSEIWP 96.2, is the principal
documentation accompanying the encoded texts. It is supplemented
by HSEIWP 96.3 and 96.4, expanding upon choices made in
implementing the syntactic and semantic encoding schemes,
respectively, as well as HSEIWP 96.1 by Kirk Lowery which
provides a general introduction to the project.

1.2  This fourth draft reflects the revised encoding scheme
employed in Version 4.1 (August 1997) of the book of Jonah.

Part I: The Logic of Syntax

2. Introduction

2.1  The central goal of the Hebrew Syntax Encoding Initiative
(HSEI) is a consistent, "theory-neutral" syntactic tagging of the
Hebrew Bible (Lowery, HSEIWP 96.1). This working paper addresses
the question of what "theory-neutral" means and how such
neutrality can be practically implemented in a syntactic tagging
scheme.

2.2  The number of actively pursued syntactic theories numbers
several dozen at present. At first blush, any consensus seems
remote indeed, and the notion of "theory-neutral" syntax at a
second remove from such a consensus. However, such a view of
theory-neutral syntax makes a number of unfounded assumptions,
and in fact abandons any unifying definition of "syntax." If we
were to turn table and start with a metatheoretical exploration
of "syntax" itself, I argue we would find a ready-made definition
of "neutrality" and even clear suggestions as to how to implement
such neutrality.

2.3  This paper is organized as follows. This Part I examines the
metatheory of syntax and the logic of syntactic explanation,
identifying the bases of a "possible syntactic explanation." Part
II extends these "possible bases" of syntactic explanation to
minimal or "neutral" tagging schemes. Part III is a general
proposal for a "theory-neutral" tagging scheme is this light;
appended are questions that are raised in this paper for further
consideration.

3. Syntactic Metatheory

3.1. The question of "theory-neutral" syntax, not surprisingly,
raises many an eyebrow. Respondents are puzzled as to how syntax
could be in any way "neutral"; they understand, quite correctly
in fact, that syntax is "theoretical" to the core. Respondents
are puzzled because it appears that we are dealing with an
obvious contradiction in terms. On this view, a morphologically
tagged database such as MORPH (Westminster Morphologically
Analyzed Machine-Readable Hebrew Bible) is already "theory-
neutral."

3.2  If we shift the question, however, to what is "syntax"?
puzzled respondents often begin to see that a metatheory of
syntax would be in some sense "neutral." Such a metatheory can in
fact be fleshed out rather uncontroversially (Moravcsik 1980,
Stockwell 1980). From such a metatheoretical starting point, we
can explore the logic of syntactic explanation. From there, a
minimal syntactic representation can be identified for each
possible basis of syntactic explanation, and a minimal tagging
scheme can be employed to satisfy these minimal requirements: in
effect a "minimal" or "theory-neutral" syntactic tagging system.
(What follows is based in part on Moravcsik 1980.)

3.3  It is not an accident that syntactic theories have a family
resemblance: there is a necessary similarity because the
metatheory of syntax imposes severe limitations on what
constitutes a syntactic description (beyond the usual scientific
requirements of explicitness, generalization, etc.)  Syntax at
this metatheoretical level can be reduced to two related
questions: 1) co-occurrence (meaningful units of sound-form) and
2) linear or temporal order of such "units." Syntax is therefore
about getting such "units" in "order."

3.4  Syntactic theories can be classified at this metatheoretical
level as well in terms of 1) what sort of facts are of primary
concern, 2) what are the goals in explaining such facts, and 3)
what means are employed in describing and explaining such facts.
This classification intersects with the problem of theory-neutral
tagging in a limited but significant way.

3.5  Along the fact dimension, theories will differ in terms of
1) size of constituent (from morpheme to clauses), 2) size of
utterances (up to sentences and beyond), 3) corpus of study, 4)
nature of utterances, and 5) relations to extrasyntactic factors
(e.g., prosody). In a fixed corpus such as the Hebrew Bible, 3)
and 4) are not questions, nor presumably will 5) enter the
discussion. As far as tagging goes, a natural limit on 2) is the
sentence, though we will want the researcher to be able to add
discourse-level annotations to the clause/sentence tag.

3.6  The question of "size of constituent" is crucial in the
determination of "theory-neutral" syntax, however.  If the depth
of constituent parsing is limited in some arbitrary manner (e.g.,
Lowery 1995), then we are presupposing particular vs.
metatheoretical syntactic statements. In this case, I believe
"theory-neutral" must at least entail exhaustive parsing down to
individual heads of all constituents. Contrast the following
analysis of the first constituent in Deut 8:1.

(a)  theory-particular (Lowery 1995: appendix, 119)

     (kol hammicwa a$er anoxi mcawwexa hayyom)

(b)  theory-neutral

     (kol ((hammicwa) (a$er ((anoxi) (mcawwe(xa)) (hayyom)))))

3.7  The range of goals does not in any way touch the question of
theory-neutral tagging; however, the matter of means or
"descriptive devices" assuredly does touch the question here. 
What is the nature of syntactic descriptions? What is a possible
syntactic description? What can syntactic facts possibly be
related to? What are such facts predictable from? Here is the
very heart of the question of theory-neutral syntax, and
fortunately the metatheoretical logic constrains the answer.

4. Bases of Syntactic Explanation

4.1  There are only three possible bases for syntactic theory:
sound, meaning, or an abstract intermediate (morphosyntax). Sound
is eliminated immediately. Meaning can be restricted to semantic
roles in predication. Morphosyntax can be broken down into three
components: 1) morphosyntactic features, 2) syntactic categories
and phrases ("parts of speech"), and 3) syntactic roles (or
"case"). Morphosyntactic features are already looked after in a
morphologically tagged database (though it would be worth
reviewing that tagging in this light).

4.2  In summary, then, the bases of syntactic explanation can be
reduced to a) constituency-dependency (phrase structure), b)
syntactic relations, c) semantic roles or d) some combination of
a)-c) (Moravcsik 1980, Stockwell 1980, Wirth 1980, reviewing
contributions to Moravcsik & Wirth 1980).

4.3  At the most general level, syntactic and semantic roles are
parasitic on formal head/constituent structure. Such role tags
can easily be superimposed on an exhaustively parsed constituent
tree. The syntactic trees can be formed, ideally, by recursively
building endocentric phrases for some syntactic category X, such
that XP --> ... X ... .

4.4  Syntactic roles or functions are given more or less
uncontroversially, e.g., subject of verb (nominative), object of
verb (accusative), etc.. The semantic roles are not so
straightforward.  Clearly a thematic hierarchy such as agent >
... > benefactive > patient, is the correct approach here.  What
is not immediately clear is how finely articulated this hierarchy
should be, nor what status should be given to non-arguments or
"adjuncts" nor how finely articulated the latter distinctions
ought to be.

PART II: THREE BASES OF SYNTACTIC EXPLANATION

5. The Skeleton of Phrase Structure

5.1  In a theory-neutral tagging, we will require that there are
as many syntactic head types as there are parts of speech. This
is crucial if we are to preserve neutrality. In my view, e.g.,
the participle is an adjectivalization, and might conveniently be
tagged <A>.  However, this is clearly not neutral; what would be
neutral is to tag a participle <PT> and allow the researcher to
configure searches that collapse <A> and <PT>. Further, to be
consistent, all participles in form must be so tagged: e.g., ro9e
"shepherd" is also <PT> rather than <N>.

5.2  The expanded set of tags is exhausted by the following.

general schema:     

<X(Y)> ... </X(Y)>,     where X and Y range over A-Z

sample list:

<A>  adjective
<AD> modifier of <V> & <A> when required;
     particles such as 'ak, 'ulay
<C>  complementizer (read subordinating conjunction)
<&>  conjunction
<I>  infinitive
<IA> so-called infinitive absolute
<P>  preposition
<PT> participle
<N>  noun
<NG> negatives (lo & al)
<V>  finite or inflected verb
<Z>  cover symbol for unassigned particles, including clitic -na'

5.3  It may be controversial if not impossible to assign a
particular tag in a given case. In this light, there should be a
declaration of all such decisions, e.g., 'ayin/'eyn "not a" as
<N>, 'ulay "maybe" as <AD>. As well, there is provision for
non-assignment using the cover symbol <Z>, with e.g., the
particle gam "even," again listing these in a taxonomic
declaration.

5.4  In the unmarked case, the syntactic head projects a
endocentric syntactic phrase. However, such a projection would be
theory-particular in the case of the logical functions <&> and
<NG>, and most likely be a problem in the case of the unassigned
<Z>. The adopted conjuction schema is:

          XP
        /    \
     XP        &P
             /    \
          &         XP

5.5  It is reasonably uncontroversial to project the clause from
C, i.e., as CP. I believe the benefits outweigh the costs in
adding an extra shell outside of the main predication. First,
this will give a consistent, unified treatment of matrix and
subordinating clauses. Second, this will give a conspicuous
parsing of the so-called "interclausal" material including the
conjunction. Third, it will permit a formal definition of the
casus pendens construction: a constituent found in the outer CP
shell.

5.6  The cost appears to be having in many cases a phrase without
an overt head. Perhaps the report engine can output 0 in this
specific case. This would in fact aid in the discourse analysis
of clause connectives (cf., e.g., Sailhamer 1990).

5.7  If the CP approach is a problem for researchers, it should
be simple to repair the situation in a global search and replace.
It would be easier to go this way than to try and insert the tags
after the fact.

5.8  The original goal was to project a nonverbal predicate
identical in formal properties to the verb phrase (VP). This was
both difficult to implement and correct, and obviously not
theory-neutral. Subject and Predicate are both treated as XPs
under government by the complementizer (C).

5.9  A theory-neutral parsing, to build on 5.4, would necessarily
be a "flat structure" vs. some sort of theory-particular
articulation (however correct it might appear to be).  Compare
two parsings of the VP in Gen 1:1.

a)   ``flat'' (PP V NP PP)

               VP
             / |  \   \
          PP   V    NP   PP


b)   ``articulated'' ((PP) (T ((NP) (t (PP))))

          TP
        /    \
     PP        T'
             /   \
          T         VP
                  /    \
               NP        V'
                       /   \
                    t         PP

5.10 There is an odd case that deserves special attention, viz.
pronominal object clitics, both on nominals and verbs. On the one
hand, we want to identify them as full NPs, in part to provide an
anchor for functional tags. On the other we do not want to
confuse them with full constituents in the study of constituent
ordering. Perhaps the output can signal the dual nature of the
objects, e.g., V-NP vs. V NP. (This is a technical matter for the
design of the report generator.)

6. Case Annotation

6.1  It is clear that some sort of notion of syntactic "case"
will have to be encoded on the constituents of the major
predication, prototypically the V heading the VP. This part of
the coding is simple and uncontroversial.

6.2  The following are the syntactic-case tags employed.

general schema:

<x> .... </x>, where x ranges over a-z

exhaustive list:

<a>  "accusative", object of V (n.b. may be used twice)
<c>  "casus pendens"
<d>  "dative", indirect object of V
<n>  "nominative", subject of V
<p>  "predicate" in verbless construction
<v>  "vocative"
<z>  cover term for unclassifiables, especially "directives"

6.3  Marginal or "peripheral" elements are simply left unmarked
for syntactic case.

6.4  A useful output, combining phrase structure and case
information, might annotate all relevant XPs.  The VP of Gen 1:1
might be represented as follows,

     PP  V  NP/n  PP/a

7. Thematic Roles

7.1  The trickiest part of the tagging scheme is the
representation of semantic functions or "thematic roles." Ideally
we want a constrained or "minimal" list with a declaration
of concise definitions for each function to ensure consistent
tagging.

7.2  The main arguments of the verb or derived verbal
(infinitive, participle, etc.) are tagged with the following
"thematic hierarchy." 

general schema:

<yz> ... </yz>, where y and z range over a-z

thematic hierarchy (borrowed from Cowper 1992: sect 3.2, 48-50),
in alphabetical order:

<ag> agent
<bn> benefactive
<ex> experiencer
<gl> goal
<in> instrument
<lc> location
<pt> patient
<pr> percept (with experiencer)
<rc> recipient (with theme)
<sr> source
<th> theme

7.3  One difficulty in implementation is the case of multiple
thematic assignments. Consider the following example (adapted
from Cowper 1992: (17c), 51).

Sue sold a     car       to Mike   for two thousand dollars.
ag
sr             th        gl/rc
gl/rc                    sr        th

7.4  The policy adopted is to conflate the roles, assigning the
thematically ``higher'' to each constituent. Therefore, the case
in 7.3 is rendered

7.5  No doubt researchers will want to distinguish semantically
among the "adjuncts" or "satellites" of the predicate as well. In
this case the list of semantic functions can be expanded to
accommodate a restricted additional list. Such a supplemental
list is taken from Lowery (1995: 1.2, 107).

<cr> circumstance
<cm> comitative
<di> direction
<du> duration
<fr> frequency
<hw> cause
<mn> manner
<pa> path
<pu> purpose
<qu> quality
<tm> time
<wh> reason

7.6  This additional list could probably be reduced, but we will
err on the side of overdifferentiation in the interests of not
biasing the case. In any case, such functions must be given clear
criteria for identification in a taxonomic declaration, e.g., a
clear difference between <pu> and <wh>.

PART III: FIRST APPROXIMATION

8. Sample Parsing

8.1  Two examples, Gen 1:1 and Deut 8:1a, will be used to
demonstrate the tagging scheme that takes all three bases of
syntactic explanation into account simultaneously, syntactic and
semantic functions parasitic on formal phrase structures. 

8.2  The text parsed is the current MORPH of the relevant
passages. The practice of conflating tags is adopted here: 
<X(Y)P:x:yz>

8.3  Genesis 1:1

>gn1:1

<CP>
<C></C>
<VP>
<PP:tm>
<P>gn1:1,1.1 B.: B.@Pp</P>
<NP>
<N>gn1:1,1.2 R")$I^YT R")$IYT@ncfs</N>
</NP>
</PP:tm>
<V>gn1:1,2.1 B.FRF^) B.R)@vqp3ms</V>
<NP:n:ag>
<N>gn1:1,3.1 ):ELOHI^YM ):ELOHIYM@ncmp</N>
</NP:n:ag>
<PP:a:pt>
<PP>
<P>gn1:1,4.1 )"^T )"T@Po</P>
<NP>
<N>
<D>gn1:1,5.1 HA H@Pa</D>
<N>gn1:1,5.2 $.FMA^YIM $FMAYIM@ncmp</N>
</N>
</NP>
</PP>
<&P>
<&>gn1:1,6.1 W: W@Pc</&>
<PP>
<P>gn1:1,6.2 )"^T )"T@Po</P>
<NP>
<N>
<D>gn1:1,7.1 HF H@Pa</D>
<N>gn1:1,7.2 )F^REC )EREC@ncbs</N>
</N>
</NP>
</PP>
</&P>
</PP:a:pt>
</VP>
</CP>

8.4  Deut 8:1a

>dt8:1

<CP>
<C></C>
<VP>
<NP:a:pt>
<N>dt8:1,1.1 K.FL- K.OL@ncmsc</N>
<NP>
<N>
<D>dt8:1,1.2 HA H@Pa</D>
<N>dt8:1,1.3 M.IC:WF^H MIC:WFH@ncfs</N>
</N>
<CP>
<C>dt8:1,2.1 ):A$E^R ):A$ER@Pr</C>
<PTP>
<NP:n>
<N>dt8:1,3.1 )FNOKI^Y )FNOKIY@pi1cs</N>
</NP:n>
<PT>dt8:1,4.1 M:CAW.</PT>
<NP:a><N>/:KF^ CWH@vpPmsX2ms</N></NP:a>
<NP:tm>
<N>
<D>dt8:1,5.1 HA H@Pa</D>
<N>dt8:1,5.2 Y.O^WM YOWM@ncms</N>
</N>
</NP:tm>
</PTP>
</CP>
</NP>
</NP:a:pt>
<V>dt8:1,6.1 T.I$:M:R^W./N $MR@vqi2mpXn</V>
<PP:pu>
<P>dt8:1,7.1 LA L@Pp</P>
<IP>
<I>dt8:1,7.2 (:A&O^WT (&H@vqc</I>
</IP>
</PP:pu>
</VP>
</CP>

8.5 Sample reports based on 8.3 (Gen 1:1) and 8.4 (Deut 8:1a,
both matrix and embedded), in that order.

8.6  Three parsed CPs (read "clauses").

CP   -->  0 VP
CP   -->  0 VP
CP   -->  C PTP

8.7  Three parsed predications, first depth.

VP   -->  PP V NP PP
VP   -->  NP V PP
PTP  -->  NP PT-NP NP

8.8  Predications annotated for syntactic case assignment.

PP  V  NP/n  PP/a
NP/a  V  PP
NP/n  PT-NP/a  NP

8.9  Predications annotated for thematic role.

PP/time  V  NP/ag  PP/pa
NP/pa  V  PP/pur
NP/ag  PT-NP/pa  NP/time

8.10 Syntactic tree of 8.3

          CP
        /    \
     C         VP
             / |  \    \
          PP   V    NP      PP
        /   \       |     /    \
     P         NP   N    PP     &P
                        /  \    / \
                       P   NP  &   PP
                                  /  \
                                 P    NP

8.11 Syntactic tree of 8.4

               CP
             /    \
          C         VP
               /    |    \
          NP        V         PP
        /    \              /    \
     N         NP        P         IP
             /    \                |
          N         CP             I
                  /    \
               C         PTP
                    /     |   \
               NP         PT       NP
                         / \       
                       PT   NP


9. Remaining Questions and Conclusion

9.1 It is not clear whether the original morphosyntactic markup
of MORPH is adequate for head-driven or word-driven approaches.
This should be revisited at some point, not necessarily before
syntactic tagging begins.

9.2  The assignment of parts of speech (syntactic head
assignment) should be reconsidered (with attention to
the category <AD>), and the treatment of unanalyzed elements
constantly revisited.

9.3  I conclude a triple tagging scheme such as the one outlined
here constitutes a "theory-neutral" syntactic tagging system.  In
addition to the morphosyntactic features already provided in
MORPH, there are three layers of tagging: 1) syntactic
constituents (head-dependent(s)), 2) syntactic function and 3)
semantic function. Both syntactic and semantic functions are
parasitic on the formal parsing.

Bibliography

Cowper, Elizabeth A.  1992. A Concise Introduction to Syntactic
Theory: The Government-Binding Approach. Chicago: University of
Chicago Press.

DeCaen, Vincent J. 1995. On the Placement and Interpretation of
the Verb in Standard Biblical Hebrew Prose. Ph.D. diss.,
University of Toronto.  Forthcoming in Copenhagen International
Seminar Series.

Lowery, Kirk E. 1995. "The Role of Semantics in the Adequacy of
Syntactic Models of Biblical Hebrew." Pp. 101-128 in Bible and
Computer: Desk and Discipline: The Impact of Computers in Bible
Studies. Proceedings of the Fourth International Colloquium of
the Association Internationale Bible et Informatique (AIBI),
Amsterdam, 15-18 August 1994.  Travaux de Linguistique
Quantitative, no. 57. Paris: Honore Champion.

Lowery, Kirk E. 1996. = HSEIWP 96.10

Matthews, P. H.  1981. Syntax.  Cambridge Textbooks in
Linguistics. Cambridge: Cambridge University Press.

Moravcsik, Edith A. 1980. "Introduction: On Syntactic
Approaches." Pp. 1-18 in Moravcsik & Wirth (1980).

Moravcsik, Edith A., and Jessica R. Wirth, eds. 1980. Current
Approaches to Syntax. Syntax and Semantics, no. 13. New York:
Academic.

Sailhamer, John H. 1990.  "A Database Approach to the Analysis of
Hebrew Narrative." Maarav 5.6 (spring): 319-335.

Stockwell, Robert P.  1980.  "Summation and Assessment of
Theories." Pp. 353-381 in Moravcsik & Wirth (1980).

Wirth, Jessica R.  1980.  "Epilogue: An Assessment." Pp. 383-386
in Moravcsik & Wirth (1980).