Lingui/stɪks/
The scientific study of human language — its sounds, structures, meanings, and life in the world — taught with a clickable IPA chart, a live syntax parser, and quantitative labs you can run.
00What Linguistics Is
Linguistics is the scientific study of language: not the mastery of many languages (that's a polyglot), but the systematic investigation of how language works in the human mind and in human societies. Linguists ask what all languages share, how they differ, how children acquire them effortlessly, and how they change over time.
The field is organized as a layered hierarchy, each level built on the one below. We'll climb it in order:
| Level | Unit of study | Core question |
|---|---|---|
| Phonetics | speech sounds (physical) | How are sounds produced and heard? |
| Phonology | sound systems (mental) | How do sounds pattern in a language? |
| Morphology | word structure | How are words built from pieces? |
| Syntax | sentence structure | How do words combine into phrases? |
| Semantics | literal meaning | What do words and sentences mean? |
| Pragmatics | meaning in use | How does context shape meaning? |
01Phonetics — The Sounds of Speech
Phonetics studies the physical reality of speech sounds (phones): how they are articulated, their acoustic properties, and how they are perceived. Because ordinary spelling is wildly inconsistent — English "ough" has six pronunciations — linguists use the International Phonetic Alphabet (IPA), where one symbol always maps to one sound.
Articulatory phonetics: describing a consonant
Every consonant is pinned down by three coordinates:
1. Voicing — do the vocal folds vibrate? /z/ is voiced; /s/ is voiceless. (Put a finger on your throat and alternate "sss"–"zzz".)
2. Place of articulation — where the airflow is constricted: lips (bilabial /p b m/), teeth (dental /θ ð/), the alveolar ridge (/t d s/), the soft palate (velar /k g/), and more.
3. Manner of articulation — how the airflow is shaped: full closure (plosive /p t k/), turbulent friction (fricative /f s ʃ/), nasal airflow (/m n ŋ/), or near-open (approximant /ɹ l j w/).
So /b/ is "a voiced bilabial plosive" and /ʃ/ (the "sh" in ship) is "a voiceless postalveolar fricative." The interactive chart in the next chapter speaks every sound.
Vowels & acoustic phonetics
Vowels are open sounds shaped by tongue position, described by height (high /i/ vs. low /ɑ/), backness (front /i/ vs. back /u/), and lip rounding. Acoustically, each vowel has characteristic resonant frequencies called formants: the first two, F1 (inversely tracks height) and F2 (tracks frontness), are enough to identify most vowels — which is why we can plot the whole vowel inventory on a 2-D chart, as the lab below does in Octave.
◓The IPA Explorer & Vowel Space
Click any consonant to see its full articulatory description and hear an English example word. (Audio uses your device's speech synthesis — example words, not the bare symbol.)
And here is the vowel space — front is left, high is top, exactly mirroring tongue position in the mouth. Click a vowel to hear its example word and read its formants.
03Phonology — The System Behind the Sounds
Phonetics catalogues every possible sound; phonology studies how a particular language organizes sounds into a mental system. The key abstraction is the phoneme: a sound that distinguishes meaning.
A phoneme is a contrastive sound unit, written in /slashes/. Its predictable physical variants are allophones, written in [brackets]. English /p/ is aspirated [pʰ] in pin but plain [p] in spin — same phoneme, two allophones in complementary distribution (each appears only where the other can't).
How do we find a language's phonemes? The minimal pair test: two words differing in exactly one sound and one meaning prove those sounds are separate phonemes.
Phonology also studies phonological rules (e.g. assimilation — the negative prefix surfaces as in-possible → im-possible before a bilabial), syllable structure (onset–nucleus–coda; the sonority sequencing principle governs which clusters are legal), and suprasegmentals like stress and tone. In Mandarin, tone is phonemic: mā (mother) vs mǎ (horse) differ only in pitch contour.
04Morphology — The Structure of Words
Morphology studies the internal structure of words, built from morphemes — the smallest units that carry meaning or grammatical function.
Free morphemes stand alone (book, run, happy). Bound morphemes must attach to something (un-, -s, -ed, -ly).
Roots carry the core meaning; affixes (prefixes, suffixes, infixes, circumfixes) modify it.
Derivational affixes create new words, often changing category: happy → happi-ness (adj → noun). Inflectional affixes mark grammatical features without changing category or core meaning: cat → cat-s, walk → walk-ed. English has only eight inflectional suffixes; derivation is open-ended.
So unbelievably = un- (negative prefix) + believe (root) + -able (adj-forming) + -ly (adverb-forming) — four morphemes, built up in layers. The analyzer below shows this for many words.
Morphological typology
Languages package morphemes differently — a spectrum:
| Type | Idea | Example |
|---|---|---|
| Isolating | one morpheme per word | Mandarin Chinese |
| Agglutinative | many morphemes, cleanly separable | Turkish, Swahili, Japanese |
| Fusional | one affix fuses several meanings | Spanish, Russian, Latin |
| Polysynthetic | whole sentences in one word | Inuktitut, Mohawk |
◓The Morpheme Analyzer
Click a word to break it into its morphemes, colour-coded by type, with a gloss of each piece and how the meaning is assembled.
06Syntax — How Words Combine
Syntax studies how words assemble into phrases and sentences according to rules every native speaker knows unconsciously. We sense that "the dog chased the cat" is well-formed and "chased dog the cat the" is not, without ever being taught why.
Constituency & phrase structure
Sentences aren't flat strings — they have hierarchical structure. Words group into constituents (phrases) that behave as units. In "the small dog," the words form a noun phrase (NP) you can replace with a single pronoun ("it") or move as a block. We capture this with phrase-structure rules:
S → NP VP
NP → Det Nom | Pronoun | NP PP
Nom → Adj Nom | N
VP → V NP | VP PP | V
PP → P NP
These few rules are recursive — NP contains PP which contains NP — so a finite grammar generates infinitely many sentences. That recursion is the engine of linguistic productivity.
Structural ambiguity
When one string fits more than one tree, it has two meanings. "I saw the man with the telescope" can mean I used a telescope (the PP attaches to the verb phrase) or the man had a telescope (the PP attaches to the noun phrase). Same words, two structures — the parser below draws both.
◓The Syntax-Tree Parser
Type a sentence using the toy lexicon below; a real chart parser builds every valid phrase-structure tree and draws it. Ambiguous sentences yield multiple trees — switch between them to see how structure encodes meaning.
08Semantics — Literal Meaning
Semantics studies meaning that is built into words and structures, independent of context. It splits into the meaning of words (lexical semantics) and how meanings combine (compositional semantics).
Lexical relations
| Relation | Meaning | Example |
|---|---|---|
| Synonymy | same meaning | big / large |
| Antonymy | opposite | hot / cold |
| Hyponymy | kind-of (subtype) | rose is a hyponym of flower |
| Meronymy | part-of | wheel is a meronym of car |
| Polysemy | one word, related senses | "head" (body / of a company) |
| Homonymy | one form, unrelated senses | "bank" (river / money) |
A key distinction is sense vs reference: "the morning star" and "the evening star" have different senses but the same referent (Venus). Sense is the concept; reference is the thing picked out in the world.
09Pragmatics — Meaning in Context
Pragmatics studies how context supplies meaning beyond the literal. "It's cold in here" can mean, on its surface, a fact about temperature — but do the work of a request to close the window.
Deixis — words whose reference depends on the utterance situation: I, you, here, now, this, tomorrow. ("Meet me here tomorrow" is useless on a note with no time or place.)
Speech acts (Austin, Searle) — utterances do things: asserting, promising, warning, christening. "I now pronounce you married" doesn't describe an act; it performs one.
Presupposition — background taken for granted: "Have you stopped smoking?" presupposes you smoked.
Grice's cooperative principle
Paul Grice observed that conversation is cooperative, governed by four maxims: Quantity (be as informative as needed, no more), Quality (be truthful), Relation (be relevant), and Manner (be clear). We convey implicatures by appearing to flout them. If asked "How was the dinner?" and you say "Well, the candles were nice," you observe Quantity's letter while implying — without stating — that the food was not. Hearers compute the implication automatically.
10Sociolinguistics — Language in Society
Sociolinguistics studies how language varies with social factors — region, class, age, gender, ethnicity, situation — and what that variation reveals.
- Dialects are full, rule-governed varieties differing in pronunciation, vocabulary, and grammar. The line between "language" and "dialect" is political, not linguistic — "a language is a dialect with an army and a navy."
- Registers are situational styles: you speak differently in a job interview than at a barbecue. Shifting between them is code-switching.
- Variation is structured. William Labov's studies showed that features like dropping post-vocalic /r/ correlate systematically with class and formality — variation follows patterns, never chaos.
- Prestige & change. "Standard" varieties carry overt prestige, but non-standard forms carry covert prestige (identity, solidarity) and often drive language change from below.
11Historical Linguistics & Cognates
Historical linguistics studies how languages change and how they are related. Change is constant and largely regular — especially sound change, which tends to apply across the board.
The comparative method uses sets of cognates — words descended from a common ancestor — to reconstruct unattested proto-languages and to build the family tree of languages (Indo-European, Sino-Tibetan, Niger-Congo, Austronesian, and more). Cognates often look similar but have diverged by regular changes: English night, German Nacht, Latin nox/noct-.
One quantitative handle on similarity is edit distance (Levenshtein distance): the minimum number of insertions, deletions, and substitutions to turn one word into another. Dialectometry and computational historical linguistics use it to measure relatedness. Try it:
12Language, Acquisition & the Mind
Psycholinguistics studies how language is processed, stored, and acquired in the brain.
First-language acquisition
Children master their native language with astonishing speed and uniformity — babbling by 6 months, first words around 1 year, two-word combinations by 2, and complex grammar by 4 or 5 — without explicit instruction and on impoverished input. They overregularize ("goed," "foots"), revealing that they extract rules rather than merely imitating. This led Chomsky to posit an innate Language Acquisition Device and the poverty of the stimulus argument: children know more than the input alone could teach.
Language & thought
Does the language you speak shape how you think? Strong linguistic determinism (language fixes thought) is rejected, but a weaker linguistic relativity (the Sapir–Whorf hypothesis) — that language can influence habitual cognition, e.g. in colour categorization or spatial reckoning — has modest experimental support.
The brain reflects this specialization: damage to Broca's area impairs fluent production while comprehension survives; damage to Wernicke's area yields fluent but meaningless speech — early evidence that grammar and meaning are partly distinct neural systems.
◓Quantitative Linguistics & Zipf's Law
Language obeys striking statistical regularities. The most famous is Zipf's law: in any large text, a word's frequency is roughly inversely proportional to its frequency rank. The most common word appears about twice as often as the second, three times as often as the third, and so on — a straight line on a log-log plot.
Paste any text below (or use the sample) to see its words ranked by frequency. Watch the steep drop-off that signals Zipf's law in action.
14The GNU Octave Linguistics Lab
Linguistics is increasingly quantitative. Below, six runnable GNU Octave programs cover Zipf's law, edit distance, vowel formants, distinctive-feature distances, letter statistics, and distributional semantics. Run them at octave-online.net or in a local Octave.
Lab 1 — Zipf's law from a word list
% A small bag of words (token list). Count, rank, and check Zipf's law.
words = {'the','cat','the','dog','the','the','cat','sat','the','dog','ran','the','cat'};
[uniq,~,idx] = unique(words);
counts = accumarray(idx, 1);
[freq, order] = sort(counts, 'descend');
rank = (1:numel(freq))';
printf('rank word freq rank*freq\n');
for i = 1:numel(freq)
printf('%3d %-7s %3d %3d\n', i, uniq{order(i)}, freq(i), i*freq(i));
end
% Zipf: freq ~ C / rank, so log(freq) is linear in log(rank).
loglog(rank, freq, 'o-'); xlabel('log rank'); ylabel('log frequency');
title("Zipf's law");
Lab 2 — Levenshtein edit distance (dynamic programming)
function d = edit_distance(a, b)
m = numel(a); n = numel(b);
D = zeros(m+1, n+1);
D(:,1) = 0:m; D(1,:) = 0:n; % base cases
for i = 2:m+1
for j = 2:n+1
cost = (a(i-1) != b(j-1)); % 0 if match, 1 if substitution
D(i,j) = min([D(i-1,j)+1, D(i,j-1)+1, D(i-1,j-1)+cost]);
end
end
d = D(m+1, n+1);
end
printf('night ~ nacht : %d\n', edit_distance('night', 'nacht'));
printf('three ~ drei : %d\n', edit_distance('three', 'drei'));
printf('water ~ wasser: %d\n', edit_distance('water', 'wasser'));
Lab 3 — The vowel space (formant plot)
% Approximate F1/F2 (Hz) for English monophthongs
labels = {'i (beet)','I (bit)','E (bet)','ae (bat)','A (father)','U (book)','u (boot)'};
F1 = [280 400 550 690 710 450 310];
F2 = [2250 1920 1770 1660 1100 1030 870];
plot(F2, F1, 'o', 'MarkerSize', 8, 'MarkerFaceColor', 'r');
set(gca, 'XDir', 'reverse', 'YDir', 'reverse'); % front=left, high=top
text(F2+30, F1, labels);
xlabel('F2 (Hz) -- backness'); ylabel('F1 (Hz) -- height');
title('English vowel space');
Lab 4 — Distinctive features & phoneme distance
% Binary distinctive features: [voiced nasal labial coronal dorsal continuant]
ph = {'p','b','m','t','d','n','k','g','s'};
M = [ 0 0 1 0 0 0; % p voiceless bilabial stop
1 0 1 0 0 0; % b
1 1 1 0 0 0; % m
0 0 0 1 0 0; % t
1 0 0 1 0 0; % d
1 1 0 1 0 0; % n
0 0 0 0 1 0; % k
1 0 0 0 1 0; % g
0 0 0 1 0 1 ]; % s voiceless coronal fricative
% Feature distance = number of differing features (Hamming).
printf('p vs b (voicing only) : %d\n', sum(M(1,:) != M(2,:)));
printf('p vs t (place only) : %d\n', sum(M(1,:) != M(4,:)));
printf('p vs n (very different): %d\n', sum(M(1,:) != M(6,:)));
% Natural classes share feature columns; minimal pairs differ in one feature.
Lab 5 — Letter bigrams & type-token ratio
text = 'the theme of these theories';
text = text(text != ' '); % strip spaces
n = numel(text);
% Count letter bigrams (pairs of adjacent letters)
bg = {};
for i = 1:n-1; bg{end+1} = text(i:i+1); end
[u,~,idx] = unique(bg);
c = accumarray(idx,1);
[c,o] = sort(c,'descend');
printf('Top bigram: "%s" (x%d)\n', u{o(1)}, c(1));
% Type-token ratio: vocabulary richness (unique words / total words)
words = strsplit('the cat the dog the cat ran');
TTR = numel(unique(words)) / numel(words);
printf('Type-token ratio = %.3f\n', TTR);
Lab 6 — Distributional semantics (cosine similarity)
% "You shall know a word by the company it keeps." (Firth)
% Context-count vectors over features [furry, barks, purrs, drives, wheels]
dog = [5 4 0 0 0];
cat = [5 0 4 0 0];
car = [0 0 0 5 4];
cosine = @(u,v) dot(u,v) / (norm(u)*norm(v));
printf('sim(dog,cat) = %.3f\n', cosine(dog,cat)); % high: both animals
printf('sim(dog,car) = %.3f\n', cosine(dog,car)); % low: unrelated
% This is the seed of modern word embeddings (word2vec, GloVe).
◓Self-Test Quiz
Twelve questions spanning every subfield. Click an answer for instant feedback and explanation.
16Reference & Further Study
| Subfield | Studies | Key concept |
|---|---|---|
| Phonetics | physical speech sounds | IPA, place/manner/voicing, formants |
| Phonology | sound systems | phoneme vs allophone, minimal pairs |
| Morphology | word structure | morphemes, inflection vs derivation |
| Syntax | sentence structure | constituency, phrase-structure rules, recursion |
| Semantics | literal meaning | sense/reference, compositionality |
| Pragmatics | meaning in context | implicature, speech acts, deixis |
| Sociolinguistics | language & society | variation, register, prestige |
| Historical | language change | sound change, cognates, comparative method |
| Psycholinguistics | language & mind | acquisition, critical period |
| Computational | language & computation | Zipf's law, parsing, embeddings |