HSEIWP 96.32

"On the Syntactic Encoding of the Hebrew Text"

Hebrew Syntax Encoding Initiative Working Papers, no 96.3
revised draft no. 2, August 1997

Vincent DeCaen
University of Toronto

* * *

Contents

1.   Introduction

Part I: Phrase Structure
2.   Syntactic Category Assignment
3.   Nonprojecting Heads
4.   Null Heads
5.   Pronominals
6.   Conjoined Structures
7.   Apposition
8.   Positioning of Embedded Clauses

Part II: Case Marking
9.   Basic Case Marking
10.  Extended Case Marking
11.  Casus Pendens

* * *

1. Introduction

1.1  The working paper HSEIWP96.2, ""Theory-Neutral" Syntactic
Tagging of Text," presents an overview of the encoding scheme and
explains the theory behind it. This paper, HSEIWP96.3, enters
into detailed consideration of the decision points in
implementing the syntactic encoding scheme in the four chapters
of Jonah. Another paper HSEIWP96.4, "On the Semantic Encoding of
the Text," considers similar decisions made in the semantic
scheme.

1.2  These three papers, together with Kirk Lowery's HSEIWP96.1
introducing the Encoding Initiative, constitute the basic package
of documentation accompanying the HSEI text of Jonah and the
projected four books of Samuel-Kings.

Part I: Phrase Structure

2. Syntactic Category Assignment

2.1  The use of the four major syntactic categories, related
three verb-like categories, and two associated functional heads
presents no difficulties.

<A>  adjective
<N>  noun
<P>  preposition
<V>  finite verb

<I>  infinitive
<IA> so-called infinitive absolute
<PT> participle

<D>  determiner (ha-)
<NG> negatives (lo and al)

2.2  <AD>.     The category "adverb" is assigned if it is
assigned thus, e.g., in the standard Brown-Driver-Briggs. A
declaration of such assignments will accompany the HSEI end
product and will be compared against the MORPH analysis. In
Jonah three such elements were tagged <AD>: 

     ak, ulay, ayin (in me'ayin).

2.3  <C>.      The category "complementizer," while in theory a
superset of the set of subordinating conjunctions, is here for
all practical purposes the set of subordinating conjunctions.
Again, a declaration of such assignments will accompany the HSEI
text. In Jonah three such elements were tagged <C>:

     a$er, ki, ha- (interrogative particle).

2.4  <&>.      In earlier versions <CJ>. The only token of the
category "conjunction" found in Jonah is wa- "and."

2.5  <Q>.      The category "quantifier" was intended to cover
miscellaneous quantifiers, but in practice cover ``numbers'' in
the book of Jonah.

2.6  <Z>.      The assignment <Z> is reserved for
"unclassifiables." A declaration of such assignments will
accompany the HSEI text. Two elements, perhaps not unrelated, are 
so designated in Jonah: the enclitic -na in "polite" imperatives;
and the vocative-related anna.

3. Nonprojecting Heads

3.1  The initial assumption was that a theory-neutral scheme
would treat all syntactic heads <X> equally, projecting an
endocentric phrase <XP>. Such a blanket treatment would be
anything but theory-neutral, however, at least for the functional
heads <D> "determiner," <NG> "negative" and <Z> "unclassifiable."
Such a treatment would introduce a specific theory of
``functional projections,'' which is clearly undesirable.

3.2  From a practical point of view, it is almost as if these
heads do not exist syntactically: they apparently affect word
order in no way; and their positioning is predictable, as a rule
operating as clitics with the associated major-category
projections.

3.3  It was decided, therefore, to "bury" these heads in
``adjoined'' structures. This strategy was adopted also with the
interrogative clitic ze (assigned <A> in Jonah 1:8). Three
samples are given.

          N                   V                   P
        /   \               /   \               /   \
     D         N         NG        V         P         A
     |                 /    \                |         |
     ha-            NG        Z              miz       -ze
                    |         |
                    al        -na

4. Null Heads

4.1  It can easily be argued that introducing the complementizer
<C> as a projecting head is not theory-neutral; indeed, some
current theories such as Head-Driven Phrase Structure Grammar
(HPSG) are explicit about not adopting this head analysis.

4.2  Nevertheless, the <C>-head analysis is adopted here. In my
view, the benefits of introducing the extra shell outweigh the
costs. First, a formal unification of main and subordinate
clauses is obtained. Second, there is some reason to believe
sorting clause types and analyzing the verbless clause will be
facilitated. Third, and perhaps most important, interclausal
connectives will receive a conspicuous parse. Fourth, it can be
used to leverage a tagging of the casus pendens construction.

4.3  If the <C>-head analysis poses in any way a problem for
researchers, and it might at the outside, then it should be a
simple matter of manipulating the tagging scheme to satisfy the
research requirements. It would be easy to do this, but might
prove difficult indeed to insert the <C> and <CP> tagging after
the fact.

4.4  The cost of inserting <C> and <CP> tags consistently as
clause markers is the ubiquitous null "complementizer." One way
to handle this case is to insert some additional "text", but on
general principles the MORPH text should be left undisturbed. The
other way to handle it is to have nothing between tags: <C></C>.
This ad hoc approach should not prove difficult for the search
and report software.

5. Pronominals

5.1  In general the independent pronouns are treated as <NP>s, as
might be expected. The double shell produced does admittedly look
redundant, <NP><N>...</N></NP>; and in fact I address such
redundancies in the phrase structure markup in HSEIWP96.6.

5.2  Morphologically bound pronouns do not in general prove a
difficulty in the markup system. For instance, the suffixed
genitive is syntactically consistent with its nonpronominal
construction.

          NP
        /    \
     N         NP
               |
               N
               |
               pro

5.3  The only real difficulty is the suffixed object on verbal
heads. A scheme similar to the burying of heads or "adjunction"
in section three is adopted that allows for the analysis of the
pronoun as a full <NP>. (It is an open question as to how this
might be dealt with by the report engine.)

          V
        /   \
     V         NP
               |
               N
               |
               pro

6. Conjoined Structures

6.1  There does not appear to be a theory-neutral way to treat
conjoined structures. I have adopted a head analysis of the
conjunction <&>, assuming that that will facilitate a
textlinguistic parsing of the text. To be consistent, that
analysis is extended throughout.

6.2  The null hypothesis is that full phrases are conjoined. The
schema employed as a default then is,

          XP
        /   \
     XP        &P
             /    \
          &         XP

6.3  In restricted cases, an analysis of head conjunction is
adopted where the conjunction of full phrases cannot be encoded.
In Jonah, this is limited to conjoined participles in 1:11 and
1:12. The analysis is consistent, then, with the related
structures in section 3.

          PT
        /   \
     PT        &
             /   \
          &         PT

7. Apposition

7.1  In limited cases, <NP>s can function as modifiers instead of
objects of a <N> head (so-called bound construction). To
eliminate the confusion between such structures, I have adopted a
schema that parallels the conjunction schema; and no doubt these
structures are related.

(a) NP2 as object             (b) NP2 in apposition

          NP1                           NP1
        /   \                         /   \
     N         NP2                 NP1       NP2

7.2  Constructions with numbers <Q> are treated as <XP>s in
apposition. Such constructions could get quite complicated as the
representative structure following demonstrates.

               QP
             /    \
          QP        &P
        /    \
     QP        NP

8. Positioning of Embedded Clauses

8.1  Two classes of problems arise in the encoding of embedded
clauses in Jonah: discourse that is syntactically a direct
object; and the ubiquitous ki-clause. The following ad hoc
policies have been adopted and consistently employed.

8.2  There is no logical difficulty with an entire poem being the
direct object of the verb to say. At the level of the analysis of
the arguments of the verb, the <CP> is positioned as object and
assigned accusative case <a>.

          VP
        /    \
     V ...     CP/a

8.3  The structure of <CP>s in an extended discourse is an
interesting question: it suggests there is a macrosyntax above
clauses that could be insightfully related to the syntax of
phrases and clauses. The policy adopted is to assign a flat
structure somewhat related to conjunction: we might call this ad
hoc analysis a "flat multiple-conjunction schema."

                CP  
             /  |   \
     ...  CP   CP   ...&P

8.4  The clause <CP> headed by the <C> ki poses a difficulty in
analysis. Generally, the <CP> can function as object or modifier
of <V>; and can move about the <VP>. There is some ambiguity,
however, as to the status of the ki-<CP> when following the verb
and its phrases: is it still within the <VP> or does it attain
the status of a full independent <CP>? Since the latter case
comes up infrequently and its bearing is more textlinguistic in
any case, and the possible construction is still consistent with
the linear ordering of the ``independent'' analysis as
<VP>-internal, I have adopted the blanket policy of analyzing
ki-<CP>s as <VP>-internal. The textlinguists are welcome to sort
out candidates for independent clause status.

Part II: Case Marking

9. Basic Case Marking

9.1   Nominative <n>.    As a rule, the subject is easily
identifiable and agrees with the verb in person, gender and
number. The case of the verbless clause presents a difficulty
only when the predicate is also a <NP>. Kirk Lowery is computer-
testing a principled analysis of such structures.

9.2   Accusative <a>.    An <XP> is assigned "object" <a> whether
or not it is a bare <NP>. Accusative case <a> is assigned also to
bare <CP>s and several <PP>s, especially those headed by the <P>
et (while understanding that the analysis of et is still
controversial).

9.3   Dative <d>.   Dative case is reserved for the garden
variety of instances generally translated into English as "to"
and "for." Such assignment can be supplemented by semantic case
roles such as benefactive <bn>, recipient <rc> (with "give"),
etc.

10. Extended Case Marking

10.1      Vocative <v>.  Four instances so tagged in Jonah: 1:14
(bis), 2:7, 4:2.

10.2      Directional <z>.    This case is used with verbs of
motion and might otherwise be collapsed with <a> in a given
analysis, especially with bare <NP>s. Such a consistent marking
permits easy comparison among <PP:goal>s, bare <NP>s and those
marked with the directional -ah.

10.3      Predicate <p>.      Originally, in verbless clauses the
predicate head <X> was allowed to project an <XP> immediately
governed by <C>, analogous to the verbal <VP> constructions.
However, this was difficult to implement because of word order
variation; and apparently violated theory-neutrality. Instead,
the major constituents of the verbless clause are assigned a flat
structure dominated by <C>, and the syntactic predicate is marked
<p>. (This eliminates the need for <o> "oblique" under previous
versions.)

11. Casus Pendens

11.1      As the name suggests, "casus pendens" is involved with
case marking, but of an unusual kind. It was thought useful to
try to mark obvious cases <c> for easy comparison. The marking
was restricted to <NP>s.

11.2      The difficulty of course is such marking presupposes a
specific theory of Hebrew syntax. The goal was to find a
consistent formal scheme that would catch the obvious cases and
that was consistent with the minimal syntactic scheme adopted for
phrase structure.

11.3      The simple formal rule adopted is that an <NP>
syntactically "outside" the <VP>, i.e., preceding the <C> in
linear order, is automatically assigned <c>.

11.4      As with the introduction of the <CP> analysis (section
4), the assignment of <c> under this formal rule is easier to
insert during the tagging and subsequently to modify, than to
insert after the fact.