Skip to main content Accessibility help
×
Hostname: page-component-848d4c4894-8bljj Total loading time: 0 Render date: 2024-06-29T11:01:14.815Z Has data issue: false hasContentIssue false

4 - Lexico-grammar

From Simple Counts to Complex Models

Published online by Cambridge University Press:  14 September 2018

Vaclav Brezina
Affiliation:
Lancaster University
Get access

Summary

This chapter focuses on the statistical analysis of lexico-grammatical features in language (such as articles, passive constructions or modal expressions). We start with a discussion of two types of approaches to lexico-grammar in corpora. The first approach uses the ‘Whole corpus’ research design and compares the frequencies of a linguistic variable (and its variants) in broadly defined subcorpora. The second approach employs the ‘Linguistic feature’ research design and carefully defines the contexts in which a particular variable can occur (i.e. its lexico-grammatical frame) and analyses factors which contribute to the occurrence of one variant of the variable as opposed to another. Following the second approach, the chapter shows how lexico-grammatical variation can be summarised using cross-tabulation and what statistical measures can be computed based on cross-tabulation summary tables. These measures range from simple percentages to the chi-squared test and logistic regression. Since logistic regression represents an advanced statistical procedure, large parts of the chapter are devoted to explaining this method and the interpretation of its output.
Type
Chapter
Information
Statistics in Corpus Linguistics
A Practical Guide
, pp. 102 - 138
Publisher: Cambridge University Press
Print publication year: 2018

Access options

Get access to the full version of this content by using one of the access options below. (Log in options will check for institutional or personal access. Content may require purchase if you do not have access.)

References

Advanced Reading

Balakrishnan, N., Voinov, V. & Nikulin, M. S. (2013). Chi-squared goodness of fit tests with applications. Waltham, MA: Academic Press.Google Scholar
Friendly, M. (2002). A brief history of the mosaic display. Journal of Computational and Graphical Statistics, 11(1), 89107.Google Scholar
Geisler, C. (2008). Statistical reanalysis of corpus data. ICAME Journal, 32, 3546.Google Scholar
Gries, S. Th. (2013). Statistics for linguistics with R: a practical introduction. Berlin: De Gruyter Mouton, pp. 247336.Google Scholar
Hosmer, D. W., Lemeshow, S. & Sturdivant, R. X. (2013). Applied logistic regression, 3rd edn. Hoboken, NJ: John Wiley & Sons.CrossRefGoogle Scholar
Osborne, J. W. (2015). Best practices in logistic regression. Thousand Oaks, CA: Sage.Google Scholar

Save book to Kindle

To save this book to your Kindle, first ensure coreplatform@cambridge.org is added to your Approved Personal Document E-mail List under your Personal Document Settings on the Manage Your Content and Devices page of your Amazon account. Then enter the ‘name’ part of your Kindle email address below. Find out more about saving to your Kindle.

Note you can select to save to either the @free.kindle.com or @kindle.com variations. ‘@free.kindle.com’ emails are free but can only be saved to your device when it is connected to wi-fi. ‘@kindle.com’ emails can be delivered even when you are not connected to wi-fi, but note that service fees apply.

Find out more about the Kindle Personal Document Service.

  • Lexico-grammar
  • Vaclav Brezina, Lancaster University
  • Book: Statistics in Corpus Linguistics
  • Online publication: 14 September 2018
  • Chapter DOI: https://doi.org/10.1017/9781316410899.005
Available formats
×

Save book to Dropbox

To save content items to your account, please confirm that you agree to abide by our usage policies. If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account. Find out more about saving content to Dropbox.

  • Lexico-grammar
  • Vaclav Brezina, Lancaster University
  • Book: Statistics in Corpus Linguistics
  • Online publication: 14 September 2018
  • Chapter DOI: https://doi.org/10.1017/9781316410899.005
Available formats
×

Save book to Google Drive

To save content items to your account, please confirm that you agree to abide by our usage policies. If this is the first time you use this feature, you will be asked to authorise Cambridge Core to connect with your account. Find out more about saving content to Google Drive.

  • Lexico-grammar
  • Vaclav Brezina, Lancaster University
  • Book: Statistics in Corpus Linguistics
  • Online publication: 14 September 2018
  • Chapter DOI: https://doi.org/10.1017/9781316410899.005
Available formats
×