Corpus Linguistics

Corpus linguistics is the study of language as expressed in samples (corpora) or "real world" text. The approach runs counter to Noam Chomsky's view that real language is riddled with performance-related errors, thus requiring careful analysis of small speech samples obtained in a highly controlled laboratory setting. Corpus Linguistics does away with Chomsky's competence/performance split, viewing that we can only ever reliably analyse language if the researcher does not interfere. In some areas there is an overlap with computational linguistics, as the latter moves towards language processing applications. This means dealing with real input data, where descriptions based on a linguist's intuition are not usually helpful. The field was established in 1967 when Henry Kucera and Nelson Francis published their classic work Computational Analysis of Present-Day American English, on the basis of the Brown Corpus, a carefully compiled selection of current American English, totalling about a million words drawn from a wide variety of sources. Kucera and Francis subjected it to a variety of computational analyses, from which they compiled a rich and variegated opus, combining elements of linguistics, language teaching, psychology, statistics, and sociology. Shortly thereafter Boston publisher Houghton-Mifflin approached Kucera to supply a million word, three-line citation base for its new American Heritage Dictionary, the first dictionary to be compiled using corpus linguistics. The AHD made the innovative step of combining prescriptive elements (how language should be used) with descriptive information (how it actually is used). Other publishers followed suit. The British publisher Collins' COBUILD dictionaries, designed for users learning English as a foreign language, were also compiled using corpus linguistics. Other corpora include The British National Corpus, a 100 million word collection of a range of spoken and written texts, created in the 1990s by Oxford and Lancaster Universities. The American National Corpus, created at Vassar, is of a similar size and was created in the early 2000s. The Brown Corpus has also spawned a number of similarly structured corpora: The Frown Corpus (early 1990s American English), the LOB Corpus (1960s British English) and the FLOB Corpus (1990s B.E).

See further

External links

 

<< PreviousWord BrowserNext >>
madness (band)
magnetic mirror
emma of normandy
herbert putnam
geosynchronous orbit
the americana
edward the confessor
open systems interconnect
halotolerance
pulsed inductive thruster
variable specific impulse magnetoplasma rocket
mongolian alphabet
specific impulse
rowrbrazzle
precognition
genetic algorithm
jupiter (god)
trusted client
harmony
harthacanute
harold harefoot
canute
vibrator
industrial sociology
kayaking
blackboard bold
type theory
melting point
cam
william butler ogden
john wentworth (mayor)
hiram college
joseph medill
carter harrison, sr.
carter harrison, jr.
nikkei 225
stephen smale
bae nimrod
jean claude killy
william hale thompson
anton cermak
jane byrne
harold washington
plasma equilibria and stability