corpus slide 1

33
What Is Corpus? By: Hafiz Haseeb Ul Hassan GC University Faisalabad, Pakistan +923007794977 [email protected] 1

Upload: linguistic-systems-inc

Post on 15-Jul-2015

94 views

Category:

Education


1 download

TRANSCRIPT

What Is Corpus?

By: Hafiz Haseeb Ul HassanGC University Faisalabad, [email protected] 1

What is a corpus?

The word corpus comes from Latin (“body”) and the plural is corpora

A corpus is a body of naturally occurring language …but rarely a random collection of text Corpora “are generally assembled with particular

purposes in mind, and are often assembled to be (informally speaking) representative of some language or text type.” (Leech 1992)

“A corpus is a collection of (1) machine-readable (2) authentic texts (including transcripts of spoken data) which is (3) sampled to be (4) representative of a particular language or language variety.” (MXT 2006: 5)

2

What is not a corpus?

A list of words is not a corpus Building blocks of language

A text archive is not a corpus A random collection of texts

A collection of citations is not a corpus A short quotation which contains a word or phrase that is the reason for

its selection A collection of quotations is not a corpus

A short selection from a text chosen on internal criteria by human beings A text is not a corpus

Intending to be read in different ways The Web is not a corpus

Its dimensions unknown, constantly changing, not designed from a linguistic perspective

Sinclair (2005)

3

What is a corpus for? A corpus is made for the study of language in a broad sense

To test existing linguistic theory and hypothesesTo generate and verify new linguistic hypothesesBeyond linguistics, to provide textual evidence

in text-based humanities and social sciences subjects

The purpose is reflected in a well-designed corpus

4

Why use corpora?

Even expert speakers have only a partial knowledge of a language A corpus can be more comprehensive and

balanced Even expert speakers tend to notice the

unusual and think of what is possible A corpus can show us what is common and

typical Even expert speakers cannot quantify

their knowledge of language A corpus can readily give us accurate

statistics

5

Why use corpora?

Even expert speakers are not always available to be consulted A corpus can be made permanently

accessible to all Even expert speakers cannot keep up

with language change A constantly updated corpus can reflect

even recent changes in the language Even expert speakers lack authority:

they can be challenged by other expert speakers A corpus can encompass the actual

language use of many expert speakers

LG3204 Corpus Linguistics 07-08 6

Benefits of corpus data Corpus data is more reliable

A corpus pools together linguistic intuitions of a range of language speakers, which offsets the potential biases in intuitions of individual speakers

Corpus data is more natural It is used in real communications instead of being invented

specifically for linguistic analysis Corpus data is contextualized

Attested language use which has already occurred in real linguistic context

Corpus data is quantitative Corpora can provide frequencies and statistics readily

Corpus data can find differences that intuitions alone cannot perceive E.g. synonyms totally, absolutely, utterly, completely,

entirely7

Corpora vs. intuitions Not necessarily antagonistic, but rather corroborate

each other and can be gainfully viewed as being complementary Armchair linguists and corpus linguists “need each other. Or

better, […] the two kinds of linguists, wherever possible, should exist in the same body.” (Fillmore 1992)

“Neither the corpus linguist of the 1950s, who rejected intuitions, nor the general linguist of the 1960s, who rejected corpus data, was able to achieve the interaction of data coverage and the insight that characterize the many successful corpus analyses of recent years.” (Leech 1991)

The key to using corpus data is to find the balance between the use of corpus data and the use of one’s intuitions

8

The corpus methodology

It is debatable whether CL is a methodology or a branch of linguistics CL goes well beyond this methodological role and has

become an independent discipline In spite of the name, CL is indeed a methodology

rather than an independent branch of linguistics in the same sense as phonetics, syntax, semantics or pragmatics These latter areas of linguistics describe, or explain,

a certain aspect of language use Corpus linguistics, in contrast, is not restricted to a

particular aspect of language - it can be employed to explore almost any area of linguistic research

9

A brief history of CL

The term corpus linguistics first appeared only in the early 1980s, but corpus-based language study has a substantial history

The history of CL can be split into two periods: before and after Chomsky

10

A brief history of CL

Before Chomsky Field linguists and linguists of the structuralist

tradition used “shoebox corpora” – shoeboxes filled with paper slips

Their methodology was essentially “corpus-based” in the sense that it was empirical and based on observed data

The work of early corpus linguistics was underpinned by two fundamental, yet flawed assumptions

The sentences of a natural language are finite. The sentences of a natural language can be collected and

enumerated. Most linguists saw the “corpus” as the only source of

linguistic evidence in the formation of linguistic theories

11

A brief history of CL

Chomsky revolution: Between 1957 and 1965 Chomsky changed the direction of linguistics from empiricism towards rationalism “Any natural corpus will be skewed. Some sentences

won’t occur because they are obvious, others because they are false, still others because they are impolite. The corpus, if natural, will be so wildly skewed that the description would be no more than a mere list.” (Chomsky 1962)

Our internal knowledge of language in human brain (competence, I-language) replaces observed data (performance, E-language)

Intuitions started to be relied on as evidence [Xiao, R. (2008) “Theory-driven corpus research: using

corpora to inform aspect theory”. In A. Lüdeling & M. Kyto (eds.) Corpus Linguistics: An International Handbook. Berlin: Mouton de Gruyter]

12

A brief history of CL Revival of CL

Corpus research was continued in a few centres (Brown, Lancaster) in the 60s-70s

The Brown University Standard Corpus of Present-day American English (Brown corpus)

Lancaster-Oslo-Bergen Corpus of BrE (LOB) The hardware still imposed some restrictions until the real

development started in the 1980s The marriage of corpora with computer technology rekindled

interest in the corpus methodology Since then, the number and size of corpora and corpus-based

studies have increased dramatically Nowadays, the corpus methodology enjoys widespread

popularity, and has opened up or foregrounded many new areas of research

13

Areas that have used corpora

Lexicography Lexical studies Grammatical studies Register/genre

analysis Language variation Contrastive analysis Translation studies Language change Language teaching

Semantics Pragmatics Stylistics Literary study Sociolinguistics Discourse analysis Forensic linguistics Computational linguistics …

14

Nature of corpus-based approach It is empirical, analysing the actual

patterns of use from natural texts It utilises a large and principled

collection of natural texts as the basis for analysis

It makes extensive use of computers for analysis, using both automatic and interactive techniques

It integrates both quantitative and qualitative analytical techniques

(Biber et al 1998: 4-5)

15

Why use computers?

Development of computer technology has revived CL

Machine-readability is a de facto attribute of modern corpora

Electronic corpora have advantages unavailable to their “shoebox” ancestors It is the use of computerized corpora, together

with computer programs which facilitate linguistic analysis, that distinguishes modern electronic corpora from early ‘drawer-cum-slip’ corpora

16

Why use computers?

Computerized corpora can be processed and manipulated rapidly at minimal cost E.g. searching, selecting, sorting and formatting

Computers can process machine-readable data accurately and consistently

Computers can avoid human bias in an analysis, thus making the result more reliable

Machine-readability allows further automatic processing to be performed on the corpus so that corpus texts can be enriched with various metadata and linguistic analyses Corpus markup and corpus annotation

17

A question for Deep Thought“Alright,” said the computer Deep Thought. “The Answer to the Great

Question...”

“Yes...!”

“Of Life, the Universe and Everything ...” said Deep Thought.

“Yes...!”

“Is...”

“Yes...!!!...?”

“Forty-two,” said Deep Thought, with infinite majesty and calm.

It was a long time before anyone spoke.

“Forty-two!” yelled someone in the audience. “Is that all you’ve got to show for seven and a half million years’ work?”

“I checked it very thoroughly,” said the computer, “and that quite definitely is the answer. I think the problem, to be quite honest with you, is that you’ve never actually known what the question is.”

Hitchhikers Guide to the Galaxy by Douglas Adams

What can we learn from this story?18

What corpora cannot do

Corpora do not provide negative evidence Cannot tell us what is possible or not possible Can show what is central and typical in language

Corpora can yield findings but rarely provide explanations for what is observed Interfacing other methodologies

The use of corpora as a methodology also defines the boundaries of any given study Importance of amenable research questions

The findings based on a particular corpus only tell us what is true in that corpus Generalisation vs. representativeness

See Unit B2 for pros and cons of corpora19

Ask corpora the right questions

Corpus linguistics as a methodology is only one of the (many) ways of doing things – “doing linguistics”

The usefulness of corpora depends upon the research question being investigated “They are invaluable for doing what they do, and what they do

not do must be done in another way.” (Hunston 2002: 20) The development of the corpus-based approach as a

tool in language studies has been compared to the invention of telescopes in astronomy If it is ridiculous to criticize a telescope for not being a

microscope, it is equally pointless to criticize the corpus-based approach for not doing what it is not intended to do

It is up to you to formulate research questions amenable to corpus-based investigation and to decide how to combine corpora with other resources

20

Testing your intuitions with BUY-BNC

http://corpus.byu.edu/bnc/

21

Most common noun in English

22

Search for [n*]

Top 10: time, people, way, years, year, work, government, day, man, world

Most common noun in adverts

23

Search for [nn*] in w-advert

Top 20: hotel, centre, time, world, holiday, day, service, year, house, facilities, range, club, bar, years, information, rooms, people, city, life, castle

Top 10 adj. in nonfiction vs. fiction

24

Top 10 in Nonfiction: aggregate, regulatory, offline, Keynesian, non-executive, macroeconomic, no-arbitrage, nationalised, short-run, pioneering

Top 10 in Fiction: Sabine, narrowed, unsmiling, flushed, clammy, navy-blue, sidelong, muttered, strangled, froggy

Distribution of phrasal verbs

25

“Talk” as a noun/verb in different registers

26

Synonyms: utter vs. sheer

27

Semantic prosody of caused

LG3204 Corpus Linguistics 07-08 28

Noun collocates of “CAUSE”:problems, damage, death, trouble, harm, concern, injury, problem, difficulties, loss, confusion, pain

data singular or plural?

29

Per million words:Singular: 776Academic: 21misc: 9.2spoken: 1.9newspaper: 1.6fiction: 0.3

Per million words:Plural: 1,035academic: 42.5misc: 8.8spoken: 0.2fiction/news: 0.1

How are women and men described?

30

reason for vs. reason to

31

Extra Practice with BNC VIEW

1) What are the top 5 modal verbs in English?

2) Is there any difference between verbs destroy, ruin, and demolish? If so, what is it?

3) Do you think the adjectives in “utterly + adjective” have anything in common? If so what is that?

4) Can we use the plural form of research as in “his researches”?

LG3204 Corpus Linguistics 07-08 32

Where to find what… BNCweb

http://bncweb.lancs.ac.uk/bncwebSignup/user/login.php BYU corpus hub: http://corpus.byu.edu/corpora.asp CQP corpus hub at BFSU (ID/pass=test):

http://124.193.83.252/cqp/ Bank of English (56M sample)

http://www.collins.co.uk/Corpus/CorpusSearch.aspx Sketch Engine: http://www.sketchengine.co.uk/ David Lee’s CL bookmarks

http://www.uow.edu.au/~dlee/CBLLinks.htm Corpus linguistics, translation, and language learning

http://sites.google.com/site/federicozanettinnet/cl-htm Corpus4u Community: www.corpus4u.org

LG3204 Corpus Linguistics 07-08 33