Exam preparation
Possible questions across all units, grouped by topic and by marks (2, 5, and 10, in the SRM pattern). These are practice questions to guide revision.
Exam preparation: question bank
2 Marks
- Define the information age and state what form most of its information takes.
- What is natural language, and on what shared assumption between speakers does it rely?
- State the great challenge of intelligent automatic text processing in one sentence.
- Name the two forces that drive automatic text processing and give the one word label for each.
- What is computational linguistics, and how is it related to applied linguistics?
5 Marks
- Explain why the computer is described as "a creature of a totally different nature" and why the word "unrestricted" makes the great challenge hard.
- Describe the competence gap and the scale gap, giving the book's example for each and one modern tool that answers each force.
- Using the comparison of humans and computers, explain why we automate text processing at all and how the strengths of each side motivate NLP.
10 Marks
- Discuss why automatic processing of natural language is important for humankind. Cover the centrality of text in the information age, the shift from mechanical to intellectual automation, the four wishes people have for a text machine and their modern realisations, and the two driving forces of competence and scale.
- Explain natural language as a communication protocol between humans and why this protocol breaks down when one party is a computer. Discuss shared context, the role of implicit information and ambiguity, and how these difficulties lead to the need for computational linguistics as the scientific foundation of intelligent NLP.
2 Marks
- Define linguistics and explain why it is described as a family of sciences rather than a single discipline.
- What is general linguistics, and what sits at the nucleus of the family of linguistic sciences?
- What does phonology study, and which NLP component does it underlie?
- Define a morph and give one example of a word built from several morphs.
- State the five core levels of language in order from the smallest sound units to speaker intention.
5 Marks
- Explain what syntax studies and describe how a sentence can be divided into subject and predicate, with the predicate covering verb and object. Name the NLP component that recovers this structure.
- Distinguish semantics from pragmatics. For each, give the core question it answers, whether it depends on context, and its NLP tie-in.
- Explain the claim that the levels of language form a layered pipeline in which each higher level assumes the ones below it. Illustrate with the mapping from each level to its NLP component.
10 Marks
- Describe the structure of general linguistics and its core levels: phonology, morphology, syntax, semantics, and pragmatics. For each level, state what it studies and the modern NLP component it underlies, and explain why higher levels are harder to automate.
- Using the utterance "It is cold in here", explain in detail the difference between the semantic reading and the pragmatic reading. Discuss why pragmatics is the hardest level for machines and how this connects to dialogue systems in NLP.
2 Marks
-
Define historical (comparative) linguistics and state why it carries two names.
-
Distinguish between synchrony and diachrony as introduced by Ferdinand de Saussure.
-
Define grammatical case and give one example of an oblique case that survives in English.
-
What is dialectology, and along which axis does it describe variation compared with diachrony?
-
Define quantitative (statistical) linguistics and name one task where it is applied.
5 Marks
-
Explain contrastive linguistics (typology) and show, with the example of articles and of grammatical case, how it classifies languages independently of their family.
-
Describe the difference between non-ergative and ergative sentence construction, and explain why ergative languages form a single typological group despite sharing no common vocabulary.
-
Compare sociolinguistics, lexicography, and psycholinguistics, stating what each studies and why it matters for computational linguistics.
10 Marks
-
Present linguistics as a family of related sciences. Place general linguistics at the centre, describe the specialised and supporting branches around it, and explain how they converge toward applied and computational linguistics.
-
Explain how comparison reveals language families. Describe the Romance, Germanic, and Slavonic families and the Indo-European super-family, explain the role of loan words, and show with the Spanish to English analogy how comparison can predict unknown forms.
2 Marks
- Define computational linguistics and state its main task.
- What is meant by "processing" of natural language in the broad sense used in this field?
- Define a word occurrence, a wordform, and a lexeme.
- State the two conditions a program must satisfy to count as linguistic software.
- In the sentence "I return the books next month, but she returns the book to me now," how many word occurrences, wordforms, and lexemes are there?
5 Marks
- Explain why a plain text editor is not linguistic software, why a simple hyphenation program only borders on it, and why a spell checker is linguistic proper. Refer to the language-dependent and large tests in each case.
- Using the example sentence, explain how the counts of 14, 13, and 10 are obtained and why intermediate totals such as 12 and 11 are linguistically inconsistent.
- Explain the relationship between word occurrences, wordforms, and lexemes, and state why a lexeme is described as a theoretical construction rather than a string.
10 Marks
- Discuss what computational linguistics is and why this course is described as "more linguistic than computational." Cover the definition and main task, the broad sense of processing, and the stability argument that favours linguistic principles over specific algorithms, giving modern examples such as parsers, information retrieval, and machine translation.
- Explain why the everyday term "word" is inadequate for a precise description of language, and develop the three replacement terms. Define word occurrence, wordform, and lexeme, illustrate each with worked examples, explain how lexemes are named and why they form dictionary entries, and relate the lexeme to the working of a spell checker or search engine.
2 Marks
- Define general linguistics in a single sentence.
- State the two conditions a program must satisfy to count as linguistic software.
- Why does a course in computational linguistics begin from notions of general linguistics rather than from programming?
- Give one reason a fundamental science such as physics or general linguistics stays useful over a long period.
- Name two linguistic resources that a language processing program relies on as integral parts.
5 Marks
- Explain why theoretical aspects of linguistics are necessary for computational linguistics, using the record of past attempts that lacked this foundation.
- Explain the analogy between physics and general linguistics, and use it to contrast a broad specialist with a narrowly specialised one.
- Apply the two tests for linguistic software to a plain text editor, a simple hyphenation program, and a spell checker, and state the verdict for each.
10 Marks
- Discuss the important role of the fundamental science in building language processing systems. Cover why programming skill alone is insufficient, what a linguistically grounded specialist can do that a coder alone cannot, and how this grounding supports work in an interdisciplinary team.
- Explain what counts as linguistic software. Define the two conditions of being language-dependent and large, justify why both are required, and classify a plain text editor, a hyphenation program, a spell checker, and a machine translation system against these conditions.
2 Marks
- Define an applied research lag and state its main cause.
- Distinguish a human-oriented description of a language from a computer-oriented one.
- Give two facts that show Spanish is a major world language.
- What is meant by multistage processing, and why is it relevant to modern pipelines?
- Why can an academic Spanish dictionary not be used directly by a computer program?
5 Marks
- Explain why most natural language processing research targets English and how this choice produces a lag in applied research for other languages such as Spanish.
- Using a comparison, explain the difference between human-oriented and computer-oriented descriptions of a language, and show why existing Spanish dictionaries and grammars fall into the first category.
- Describe the state of research and commercial effort on Spanish around 2004, naming the main groups and the Microsoft effort, and comment on why the level of activity was considered low.
10 Marks
- Assess the resource picture for computational processing of Spanish around 2004. Discuss the available dictionaries and grammar manuals, explain why they were not usable as direct sources for programs, identify the research teams and commercial developments working on the language, and evaluate the claim that the resource gap was as much an opportunity as a deficit.
- "A description that is excellent for a human learner is not automatically usable by a computer." Explain this statement in the context of applied research on Spanish. Define the key terms, illustrate the point with dictionaries and grammar manuals, and argue why building processing tools for under-served languages remains an important research direction today.
2 Marks
- Define a linguistic theory.
- What is the structuralist approach in general linguistics?
- What are the immediate constituents of a sentence?
- Name Chomsky's two main contributions to the description of language.
- What is a context-free grammar, and where does it sit in the Chomsky hierarchy?
5 Marks
- Explain why general linguistics allows several approaches to coexist, using a comparison with the exclusive theories of physics.
- Describe the constituency, or phrase structure, approach and show how a sentence is split into immediate constituents down to single words.
- Contrast European and American structuralism, and explain the role Bloomfield gave to word order in revealing sentence structure.
10 Marks
- Trace the historical development of grammar theory from Saussure's structuralism through the European and American schools to Bloomfield's use of word order and the constituency approach. Explain what structuralism treats language as, why the American school stressed objective and observable facts, and how sentences were split into immediate constituents.
- Explain Chomsky's initial contribution to linguistics in the 1950s. Describe the mathematical nucleus of generative grammars and formal languages, the attempt to describe real languages, and how the phrase structure idea was formalised as context-free grammars that became the basic tool for describing natural languages. Include Chomsky's own caution about treating the CFG approach as the only possible one.
2 Marks
- Define a context-free grammar as used to generate simple sentences, and name its initial symbol.
- Distinguish a non-terminal symbol from a terminal symbol, giving one example of each from the grammar.
- What is a production rule, and what is the relation between the symbols on its right side and the entity on its left?
- What does the metasymbol
|mean in the terminal rules? Illustrate with the rule for N. - Define the constituency tree of a sentence.
5 Marks
- List the four production rules of the grammar and explain, rule by rule, how they combine to derive the string D N V D N.
- Using the terminal rules, show how the sentences Mary sees the building and the student sings a song are generated, and identify the category string behind each.
- The grammar produces both the song sings a student and a songs. Explain why each is produced, say which is meaningless and which is ungrammatical, and explain why fully eliminating meaningless sentences is difficult, referring to the role of the initial symbol S.
10 Marks
- Take the sentence the student sings a song. Give its full derivation from S as a step-by-step table showing the rule used and the string after each step, draw its constituency tree, and write the equivalent labelled bracketed nested structure, explaining how the tree, the brackets, and the sequence of rule applications correspond.
- Present the complete simple context-free grammar: its non-terminal symbols and their grammatical categories, its production rules, and its terminal rules with the metasymbol for alternatives. Then explain how the grammar controls form but not meaning, using generated grammatical, meaningless, and ungrammatical sentences to illustrate how sense accumulates constituent by constituent.
2 Marks
- Define a transformational grammar.
- Distinguish between deep structure and surface structure with one example each.
- State what a transformation rule is and give one example of the operation it performs.
- Why does the sentence "Does John see Mary?" not allow a clean nested representation?
- Name the three classic transformations that Chomsky set out to explain.
5 Marks
- Explain why simple context-free grammars proved insufficient for describing natural language, and state how transformational grammars extend rather than replace them.
- Describe Chomsky's two-part proposal for handling a difficult sentence, naming both parts and explaining the role of each.
- Explain the analogy that a set of transformational rules functions like a program, identifying the input, the processing, and the output.
10 Marks
- Work through the transformation of the affirmative sentence "John sees Mary." into the interrogative "Does John see Mary?", listing every ordered step, identifying which step violates the nested structure, and relating the process to the deep and surface structures.
- Assess transformational grammars as a theory: describe what a transformational grammar is, the operations it uses, the descriptive problems it solved, and the cost it paid in mathematical elegance and explanatory power. Support your assessment with reference to the deep-to-surface mapping.
2 Marks
-
Define valency in linguistics, and state the difference between a valency and a valency filler.
-
What is subcategorization of verbs?
-
State the three actants of the verb give and the role each one plays.
-
Distinguish the syntactic aspect of a valency from its semantic aspect.
-
What is a semantic case, and who introduced the notion?
5 Marks
-
Explain, with the verb give and the sentence John gave a book to Mary, what a subcategorization frame is, what information it keeps, and what it loses compared with the semantic pattern give thing-given to receiver.
-
Take the sentence John gave a book to Mary yesterday in the library. Sort every piece into actants and circonstants, applying both the syntactic criterion and the semantic criterion, and explain why the subject is usually left out of subcategorization frames.
-
Using John gave a book to Mary and John gave Mary a book, show why one verb with a single meaning can require several subcategorization frames, and explain why this breaks the classification into non-intersecting subcategories.
10 Marks
-
Trace how linguistic research after Chomsky stayed within the phrase-structure and context-free mainstream while adding new descriptive methods. In your answer define subcategorization, actant, and valency, use the chemistry analogy, and work through the verb give to show how valencies are expressed by word order and preposition.
-
Explain why a purely generative grammar can assign syntactic roles but cannot give the semantic interpretation of a sentence. Use the pair John gave a book to Mary and John gave Mary a book to show that the syntactic label on Mary changes while her role does not, then describe how Fillmore's semantic cases were introduced to close this gap by connecting subcategorization frames to meaning.
2 Marks
-
Define a feature in the constraint-based description of grammar, and give one example.
-
What is a constraint in a grammar rule? State what it requires of the constituents.
-
Define agreement and name two grammatical properties on which constituents commonly agree.
-
Why does the rule NP to D N overgenerate in English? Give one incorrect string it admits.
-
What is Generalized Phrase Structure Grammar (GPSG), and what use did it make of features?
5 Marks
-
Write the featured rule for an English noun phrase that enforces number agreement, and explain how the shared variable Num causes it to reject *this books and *a books while accepting this book and these books.
-
Describe the three equivalent notations for stating a constraint: spelled-out feature variables, per-feature numbered indices, and a single index for a whole feature set. Explain the indexing rule by which a number first names a value and is then reused.
-
Explain how writing grammar rules with features reduces the total number of rules. Use the case of three persons and two numbers to show how one generalized rule replaces six specific rules.
10 Marks
-
Explain the constraint-based approach to grammar in full. Cover why plain phrase-structure rules overgenerate, the definitions of feature, constraint, and agreement, the feature-annotated rule for determiner and noun agreement, and the subject and predicate agreement rule S to NP(Pers, Num) VP(Pers, Num) with the examples she sings and they sing. Conclude by explaining how this approach reduces the number of rules and how it leads toward unification.
-
Compare the constraint-based approach with the plain phrase-structure and transformational descriptions that precede it. Discuss what each contributes, how constraints separate generation from filtering, how the numbered-index or reentrancy idea works, and why matching feature sets is a general operation rather than a per-rule device. Relate the approach to GPSG and to modern feature-based parsers.
2 Marks
- Define Head-Driven Phrase Structure Grammar (HPSG) and name the earlier formalism it succeeds.
- What is the head (head daughter) of a constituent, and how is it marked in a rule?
- State the head principle in one sentence.
- Define unification as used in computational linguistics.
- When is a feature set said to be underspecified? Give one English example.
5 Marks
- State the head principle and explain, using the rules
S → NP HVPandNP → D HN, how features flow from the head daughter to the mother constituent. - Explain the compatibility condition for unification. Show why the feature sets of "this" and "boy" unify, and why those of "this" and "boys" do not.
- Describe how HPSG attaches meaning to words. Contrast this with the way early generative grammars handled semantic interpretation, and note the role of transformations.
10 Marks
- Work through the unification of "these sheep" and of the verb "was" in "The boy was late". For each case, identify the feature sets, show which values are underspecified or ambiguous, give the unified result, and explain how the head principle then fixes the feature of the enclosing phrase or sentence.
- Compare the head principle and unification as the two core mechanisms of HPSG. Cover their roles, the direction in which information moves, their conditions, and their effects, and explain how a full parse uses both together to keep agreement consistent while filtering out incorrect analyses.
2 Marks
- Define Meaning-Text Theory and state where and when it was proposed.
- List the five representation levels of the MTT transformer in order from meaning to text.
- What is a valency in linguistics? Give one example of a word with three valencies.
- Which symbols label semantic valencies and which label syntactic valencies in a government pattern?
- Why is a language in MTT described as a reversible transformer rather than a generator?
5 Marks
- Explain the notion of the multistage transformer in MTT. Describe what a representation is, how the levels relate to one another, and how the model differs from the Chomskian notion of transformation.
- Define a government pattern and describe the three parts of a verb's government pattern. Use the verb explain to illustrate how the parts fit together.
- Distinguish semantic valencies from syntactic valencies. Explain, with an example such as the English verb give, why the correspondence between them is not always one to one and how multiple realisation options are recorded.
10 Marks
- Present the government pattern of the verb explain in full. Read each part of the table, explain the meaning of the symbols X, Y, Z and 1, 2, 3, the part-of-speech labels, and the role of the preposition to, and then discuss how the pattern changes for verbs that allow several options per slot and restricted word orders, using the verb give as the contrasting case.
- Compare government patterns of Meaning-Text Theory with the subcategorization frames of generative grammar. Cover the tradition each belongs to, the starting point of each, what each records, and their scope across parts of speech, and evaluate the claim that government patterns are all-sufficient for language description.
2 Marks
- Define the Meaning-Text Theory (MTT) and name the linguist who proposed it.
- What is a dependency link? State its direction and give one example.
- Define a dependency tree, naming its nodes, its edges, and its root.
- Define a semantic link and say how it differs from a syntactic link.
- In whose work did dependency trees first appear as an object of linguistic research, and in which decade?
5 Marks
- Contrast the dependency approach with the constituency approach. Cover the nodes and edges of each kind of tree, the founders and recent developments of each tradition, and how each treats word order.
- Explain why dependency trees describe word order and agreement more easily than constituency trees, and why they can represent disrupted, non-projective constructions that nested phrase structures cannot.
- Explain why the dependency approach needs no Chomskian transformations for English interrogatives, and describe the terminology clash in which "deep structure" and "transformation" appear in both the generative and MTT traditions with different meanings.
10 Marks
- Trace how the MTT moves from the surface syntactic level through the deep syntactic level to the semantic representation. Explain the deletions, insertions, and inversions involved, illustrate a word that disappears using the auxiliary have in they have asked, and illustrate a restored element using the NAME node inserted into his son John.
- Draw and explain the dependency tree of the sentence The clever student sings a song, identifying the root and the dependents, then draw the constituency tree of the same sentence and state three differences between the two. Finally explain meaning-preserving paraphrase using John gave me help and John helped me, and say why these deep-level rules are independent of the logic-based rules at the semantic level.
2 Marks
- Define an applied linguistic system and give one example.
- What is text preparation in the broad sense, and which class of software commonly bundles it?
- Define information retrieval and state the goal of relevance ranking.
- Distinguish automatic translation from a natural language interface in one sentence each.
- What do optical character recognition and speech recognition each convert, and why are they called boundary tasks?
5 Marks
- List the main classes of applied linguistic systems and give one concrete example of each.
- Describe the five subtasks of text preparation in order of increasing linguistic difficulty, and explain why spell checkers are more mature than style checkers.
- Compare present products with prospective products of computational linguistics using status, maturity, and scope, and place three example applications into the correct column.
10 Marks
- Explain the full classification of applied linguistic systems. For each class, give a definition, a familiar example, and a note on its current maturity, and explain why the field is surveyed by its products before its methods.
- Discuss the meaning-oriented classes of applied systems, namely factual data extraction, text generation, and natural language understanding, together with the recognition tasks that feed them. Describe the input and output of each, how they connect in a pipeline, and where modern NLP systems fit into this scheme.
2 Marks
- Define automatic hyphenation and state its main purpose.
- What linguistic knowledge does a simple hyphenation program need, and give one Spanish example of an inseparable combination.
- Define spell checking and explain what "out of context" means in its definition.
- Distinguish a typographic error from a spelling error, with one example of each.
- What is an orthographic corrector, and how does it differ from a program that only detects errors?
5 Marks
- Explain where a word may and may not be split during hyphenation, using the examples teach-er, wash-ing, and graph-ic to show correct breaks and teac-her, was-hing, and grap-hic to show illegal ones.
- Describe the two stages of spell checking, detection and correction, and explain how a checker proposes candidate corrections, using the strings greit and systema as examples.
- Explain why a simple rule-based hyphenation program is adequate for Spanish but a dictionary-based program is needed for English, referring to morphemic structure and word origin.
10 Marks
- Compare automatic hyphenation and spell checking in full. For each, give a definition, its purpose, whether it uses context, the linguistic knowledge it requires, and the additional knowledge needed for higher quality. Explain why spell checking needs far more information than hyphenation and how morphology helps keep its dictionary compact.
- Discuss the limits of spell checkers that judge words out of context. Explain why real-word errors such as then for than and form for from go undetected, why corrections such as teached to taught or childs to children are missed, and what kind of more powerful tools would be required to overcome each limit.
2 Marks
- Define grammar checking and state how it differs from spell checking.
- What is agreement in a language? Give one example of an agreement error.
- Define style checking and state the kind of error it targets.
- Why can a plain spell checker not detect a real-word error, as in their typed for there?
- Name three simple statistics that commercial style checkers compute to assess a text.
5 Marks
- Explain, with examples, three types of grammatical error that a grammar checker must handle, and state why distant agreement is hard to detect.
- Explain why simple narrow-context checks fail for grammar checking, illustrating both a missed error and a false alarm.
- Describe the linguistic knowledge a capable style checker would need, and contrast it with what commercial style checkers actually do.
10 Marks
- Discuss grammar checking as a text-preparation task. Define it, classify the main types of grammatical error with examples, explain why full syntactic parsing is required for a reliable solution, and justify why the final correction decision must be left to the author.
- Compare grammar checking and style checking in detail. For each, give a definition, the kind of error it targets, the linguistic knowledge it requires, concrete examples, and its current commercial maturity, and explain why deep style assessment is still considered a task for the future.
2 Marks
- Define a reference to a word and state what it is used for.
- What is a collocation? Give one example that shows why word-by-word translation fails.
- Define recall and precision in an information retrieval system.
- What are keywords in a bibliographic database, and how is a simple query formed from them?
- What is EuroWordNet, and what is a synset?
5 Marks
- Explain the "m of n" threshold used to rank retrieved documents, and describe what happens at the two extremes m = n and m = 1.
- Describe the internal operations a dictionary of word combinations such as CrossLexica performs when a user enters a word, and identify the morphologic and syntactic issues involved.
- Explain automatic indexing and thesaurus enrichment in an IRS, and state why enrichment usually increases average recall. Use the "conjugation" and "morphology" example.
10 Marks
- Discuss references to words and word combinations as a product of computational linguistics. Cover the purpose of such references, the two kinds of linguistic tool, why the term thesaurus is used loosely, the notion of a collocation, and the role of resources such as CrossLexica, WordNet, and EuroWordNet, including the main semantic relations recorded between synsets.
- Explain how an information retrieval system works and how its quality is measured. Cover keyword sets and logical queries, ranking by relevance, precision and recall and their trade-off, automatic abstracting, and the main factor that limits accurate retrieval today.
2 Marks
- Define topical summarization in one sentence.
- What is morphological normalization, and why does Classifier perform it before counting?
- What are stop words, and how does TextAnalyst use them?
- State two practical tasks that automatic topic finding supports.
- When are two words considered related in TextAnalyst?
5 Marks
- Distinguish topical summarization from full summarization by contents. State what each conveys, the form of its output, and the question each answers.
- Explain the two kinds of linguistic information used by the Classifier system, and describe how it moves from a single word to a main topic.
- Explain how TextAnalyst assigns importance to words. Include the roles of frequency, the relationship network, and the dynamic neural network algorithm, and name the two outputs the ranked words produce.
10 Marks
- Compare the Classifier and TextAnalyst systems as approaches to topic finding. Cover the basis of each system, whether each needs a dictionary, the main output of each, and explain clearly why Classifier is called knowledge-based while TextAnalyst is not.
- Describe the Classifier pipeline in full, from raw text to a chosen main topic, using the example chain from a single word up to a broad science topic. Then discuss one known limitation of the method, such as skipped pronouns or zero subjects, and explain how it makes the gathered statistics incorrect.
2 Marks
- Define machine translation and state why the task is important.
- What is word-by-word translation, and why did the early hope in it fail?
- What is an interlingua, and what did the interlingua approach aim to achieve?
- Name one genre for which commercial translators work well and one for which they fail.
- Define word sense disambiguation and give the Spanish word used in the book to illustrate it.
5 Marks
- Explain word sense disambiguation as a difficulty in machine translation, using the word "bank" as an example, and state what is needed to resolve it correctly.
- Explain, with the book's examples, why recovering implicit information is a difficulty in machine translation, and why the target language forces choices the source text never marked.
- Assess the quality of current translation software: describe what it can and cannot do, why post-editing is often costly, and why translation quality differs by direction.
10 Marks
- Trace the two historical approaches to machine translation. Define word-by-word translation and the interlingua approach, explain why each proved inadequate, and describe the present state of commercial translation quality including the genres where it succeeds and fails.
- Discuss the two central difficulties of machine translation in detail. Define word sense disambiguation and implicit-information recovery, give a concrete example of each from the book, explain why both require deep linguistic analysis rather than dictionary lookup, and describe how directionality and research such as statistical methods and the Meaning to Text model relate to these difficulties.
2 Marks
- Define a natural language interface to a database and state the task it performs.
- What is a sublanguage, and why is the sublanguage of a database interface usually limited?
- State the meaning of the elliptical reply "And narrow?" when it follows the question "Are there wide high-resolution matrix printers in the store?"
- Define extraction of factual data from texts (information extraction).
- Which branches of linguistics are most demanding for natural language interface work, and which are less demanding?
5 Marks
- Describe the structure of an application system fronted by a natural language interface, explaining how a user question becomes an answer.
- Explain why analysis of an interface sublanguage is simpler than machine translation, and why formal query languages such as SQL reduce the pressure for a natural language interface.
- Explain the elliptical question problem in dialogue and describe how analysis through synthesis is used to resolve it.
10 Marks
- Explain what a natural language interface does, the demands it places on the different branches of linguistics, the reasons such systems succeed only for specialised sublanguages, and the role speech recognition could play. Use the printer store dialogue to illustrate the dialogue problem.
- Describe the extraction of factual data from on-line texts. Explain what a factographic database is, what parameters may be extracted, why the task is worth automating rather than using human summarizers, and why it remains a difficult and largely unsolved problem.