Department of English and Communication: A Deep-Dive Dossier on Professional Communication, Translation and Corpus Linguistics
Module 01: Academics · Sub-file: Institutional history of the Department of English and Communication (focus on the corpus-linguistics tradition) An English department that doesn't study Shakespeare, but instead records the real English of Hong Kong bankers in meetings, engineers writing tenders, and the ICAC drafting corruption-prevention guidelines — that's not something you see every day at a Chinese university. This piece will not repeat the department's current overview, QS rankings, or personnel appointments (see the Faculty of Humanities deep-dive dossier (including the Department of English and Communication chapter)). Instead, it takes an institutional-history approach to answer a more fundamental question: how this department came to make "corpus linguistics" its foundation. Sources draw mainly on PolyU's official pages, the Research Centre for Professional Communication in English (RCPCE) corpus pages, and academic press catalogues. For the department's place within PolyU's broader faculty and department structure, see Faculty and Department Overview.
The one-sentence verdict: The Department of English and Communication at PolyU is built on "language excellence in professional contexts." Since the HKCSE in the 1990s, it has assembled a family of industry corpora spanning finance (circa 7.34 million words), engineering, and anti-corruption, and originated the concgram methodology.
Why is it called "English and Communication" rather than just "English"?
The name of PolyU's Department of English and Communication (ENGL) contains the clue to its divergence from a traditional English department. Most universities' English departments centre on English and American literature, with curricula built around poetry, fiction, and literary criticism. PolyU's version, by contrast, is anchored in "Communication" — how language is used in real professional settings, rather than how it achieves beauty in literary classics.
This orientation is written into the department's official mission. According to the department's official overview page※, the department's primary goal is to provide "first class English language education with a focus on the skills and expertise needed to succeed in the global and the local workplace" — in other words, first-class English education centred on "the skills and expertise needed to succeed in the global and the local workplace." The word "workplace" appears twice in a single mission statement, nailing the department's academic identity firmly to the applied side of the fence.
At the institutional level, this orientation is not mere rhetoric but built into the structure. The department organises its research into five major areas, with "Language and Professional Communication" at the top — covering lexicogrammar, discourse, genre, and pragmatics in professional contexts, as well as organisational communication, intercultural workplace communication, and English for Specific Purposes (ESP). Literature and language arts have not vanished, but they have been compressed into an optional specialism within a taught master's programme, rather than remaining the department's centre of gravity.
To understand where this "workplace-first" identity came from, you have to go back to the 1990s — back to a corpus that set out to capture, verbatim, the English that Hong Kong people actually speak. That corpus was not just a body of data; it was the starting point of this department's methodological identity: not to infer how language should be from textbooks, but to observe how language actually is from real data.
It all began with a corpus that could "listen": how was the HKCSE built?
The Hong Kong Corpus of Spoken English (HKCSE) is the foundational project of this tradition. According to the catalogue page※ of academic publisher John Benjamins, the corpus was compiled between the mid-1990s and the early 2000s by three scholars — Winnie Cheng, Chris Greaves, and Martin Warren — within what was then the "Department of English" at PolyU. The department's current name, "English and Communication," had not yet been adopted; the corpus predates the name.
The HKCSE contains approximately 1 million words, divided into four equal sub-corpora: academic, conversational, business, and public discourse. All recordings are of adults whose native language is typically Cantonese or English — a design that deliberately preserves Hong Kong's character as an intercultural context. What was recorded was not "standard British English" but the English that Hong Kong people actually converse in, in real situations. The corpus therefore documents a site of interlanguage and intercultural communication, not a model of normative English.
The HKCSE's most unusual feature is that it was transcribed both orthographically and prosodically — capturing not just what words were said, but how they were said: where stress falls, pitch rises and falls, and the movement of intonation. According to the catalogue page, the 2008 monograph A Corpus-driven Study of Discourse Intonation (Studies in Corpus Linguistics, vol. 32) was the "first" work to systematically apply linguist David Brazil's four Discourse Intonation systems — prominence, tone, key, and termination — to a corpus of authentic natural spoken language.
This meant the PolyU team was working not with quiet written texts, but with noisy, overlapping, accented, and inflected real speech. To search this intonationally-annotated corpus, they even had to build their own tools. It was precisely this path — "real data first, new methods forced into existence" — that pushed the department toward its next original contribution: the concgram.
What is a concgram? How did PolyU scholars rewrite the definition of "collocation"?
In corpus linguistics, counting "n-grams" (n consecutive words, such as "in order to") is an old practice. But in real language, a set of words that frequently co-occur are often not adjacent and do not appear in a fixed order. For example, between "play" and "role" you can insert "an important," "a key," "a decisive" — and the pair can even be reversed, as in "the role played by." Traditional n-grams cannot capture these loose yet real collocations.
The PolyU team's solution is called the "concgram." According to From n-gram to skipgram to concgram※, published by John Benjamins in the International Journal of Corpus Linguistics (vol. 11, issue 4), a concgram is defined as a collocation generated by the association of two or more words, encompassing all permutations of constituency variation and positional variation. In other words, it treats any set of words that repeatedly co-occurs — regardless of how many words intervene, regardless of their order — as the same collocation.
The concept was subsequently built into software. According to John Benjamins' software catalogue※, ConcGram 1.0: A Phraseological Search Engine, developed by Chris Greaves, was officially published in 2009 as the first entry in the "Studies in Corpus Linguistics: Software" series. The engine automatically detects co-occurrences exhibiting both constituency and positional variation, making it, in the publisher's words, "uniquely suited to revealing the complete collocational profile of a text or corpus."
The methodology's weight has also been endorsed by peer-reviewed journals. "Uncovering the extent of the phraseological tendency: towards a systematic analysis of concgrams," co-authored by Cheng, Greaves, Sinclair, and Warren, appeared in 2009 in the top-tier journal Applied Linguistics (vol. 30, issue 2); the related research received funding from Hong Kong's Research Grants Council (RGC) (project reference PolyU 5459/08H). Turning a self-coined term into an SSCI top-journal publication, a commercially-released piece of software, and a competitive research grant — that is the complete chain of "institutionalising" a methodology, and every link in the chain lands in this department.
What is worth spelling out is what this method challenged. The traditional view of language tends to see sentences as "grammar provides the skeleton, vocabulary fills the slots." Corpus linguistics, by contrast, discovered that language is heavily reliant on strings of semi-fixed word chunks — what the late corpus-linguistics pioneer John Sinclair called the "idiom principle." The phrase "phraseological tendency" in the title of the paper above is a direct continuation of this thesis: language is far more "prefabricated" than grammar books imagine. The concgram's contribution is to attach to this abstract thesis a tool that can exhaustively enumerate and quantify — it turns "how prefabricated is language" from a philosophical assertion into a measurable corpus fact. The fact that Sinclair himself appears on the paper's author list also connected the PolyU team to the mainstream lineage of international corpus linguistics.
This methodology maps directly onto industry corpora as well. According to a publicly available analysis report from RCPCE (Mike Report※), researchers used concgrams to examine the collocational networks of terms such as "attributable" and "shareholders" in the Hong Kong financial services corpus — training an abstract method directly on real financial texts to see how words in professional English actually cluster. Method and corpus converge here: self-built tools, applied to self-built industry data, answering real questions about "professional communication."
Once the HKCSE was recorded, what did PolyU scholars read out of it?
The true value of a corpus lies in how much research can grow out of it. After its completion, the Hong Kong Corpus of Spoken English (HKCSE) became a rich seam that the department mined for two decades — proof that "recording real language" was not the endpoint but the starting point of a series of discoveries.
The first seam is "vague language." According to "The Use of Vague Language Across Spoken Genres in an Intercultural Hong Kong Corpus" (Springer chapter page※), PolyU scholars took representative samples from the HKCSE's four sub-corpora — academic, business, conversation, and public — and analysed how Hong Kong speakers use expressions such as "and so on," "or something," and "about" in intercultural settings: expressions that seem imprecise yet actually perform social functions. Vagueness is not a linguistic deficiency but a strategy for lubricating intercultural communication — a finding only achievable through real data. Earlier related research traces back to "The Use of Vague Language in Intercultural Conversations in Hong Kong," published in English World-Wide (vol. 22, issue 1).
The second seam is the pragmatic function of individual words. In work belonging to the same lineage as the International Journal of Corpus Linguistics (vol. 6, issue 2) study※, the PolyU team specifically tracked the uses of the word "actually" in intercultural conversational corpus data — how one inconspicuous adverb is used to mark topic shifts, corrections, and to soften disagreement. Turning one small word into a full paper is precisely the "great things from small details" ethos of corpus linguistics.
The third seam pushed the HKCSE toward speech act annotation. According to RCPCE's HKCSE-SAC page※, researchers built on the HKCSE to annotate 45,372 English spoken discourse units and 69 speech act types, producing a searchable speech-act sub-corpus; the corpus was compiled by Dr. Seto Wood-hung Andy as part of his doctoral thesis. Its quantitative results show that although the six genres — meetings, telephone calls, informal office conversations, service encounters, question-and-answer sessions, and interviews — have very different interactional contexts, the number and types of distinct speech acts, as well as their frequencies and co-occurrence patterns, display cross-genre similarities. Real professional speech is more alike in its deep structure than it appears on the surface.
Taken together, these three seams show one thing: the Department of English and Communication did not build a corpus and shelve it. Around the HKCSE, it grew an entire research community spanning vague language, pragmatics, and speech acts. One corpus, two decades of mining, dozens of papers and several monographs — that is the real weight of the "corpus tradition," and the core asset that distinguishes this department from an ordinary English-teaching unit.
How large is the RCPCE's "family of industry corpora"?
The foundation laid by the HKCSE and the concgram eventually grew into a dedicated research unit: the Research Centre for Professional Communication in English (RCPCE). It was no longer satisfied with a single general spoken corpus; instead, it set out to collect, category by category, the real professional texts of Hong Kong's industries — turning "language in professional contexts" from a slogan into a searchable database cluster.
According to RCPCE's industry corpus portal page※, the Centre collects authentic texts, discourse, and genres from different professional communities and contexts in Hong Kong, building a series of profession-specific corpora that are open to the public and searchable online. What these corpora share is that their data comes from English produced by engineers, financial professionals, anti-corruption officers and others in their actual work — not from textbooks or news. Below is the main membership of this "corpus family," listed by documented size:
| Corpus | Contents | Size (as specified) | Source page |
|---|---|---|---|
| Hong Kong Corpus of Spoken English (HKCSE, prosodic version) | Real spoken English across academic/conversation/business/public domains, with intonation transcription | approx. 1 million words | SCL 32 catalogue※ |
| Hong Kong Financial Services Corpus (HKFSC) | 25 text types from the financial services industry | 7,341,937 words | HKFSC page※ |
| Hong Kong Engineering Corpus (HKEC) | Authentic English texts from the engineering profession | approx. 9.2 million words | RCPCE portal※ |
| Surveying and Construction Engineering Corpus | Surveying and construction engineering texts | approx. 5.7 million words | RCPCE portal※ |
| Corpus of Research Articles (CRA) | Research papers across 39 disciplines, with sub-corpora by discipline/field/section | Multi-disciplinary (see below) | RCPCE portal※ |
| Hong Kong Corpus of Corruption Prevention (HKCCP) | ICAC corruption-prevention guidelines and related legislation | 567,599 words | HKCCP page※ |
Note on figures: The word counts in the table above are the collection sizes recorded on each corpus's official page and in publication catalogues, measured as cumulative tokens at the time the corpus was completed. For the engineering and surveying corpora, word counts are rounded to whole orders of magnitude (approx. 9.2 million, approx. 5.7 million) per RCPCE's public descriptions; the precise figures as latest published on the Centre's pages take precedence.
The numbers alone reveal the centre of gravity: finance (approx. 7.34 million words) and engineering (approx. 9.2 million words) are the two heaviest blocks, corresponding neatly to Hong Kong's two pillar industries as an international financial centre and a construction-dense city. The distribution of the corpus family is itself a sketch-map of Hong Kong's professional English landscape — the department's choice of which industries to record reflects its judgement about which languages this city runs on.
The Corpus of Research Articles (CRA) extends the reach into academic writing itself. According to RCPCE's published information, the CRA collects research papers across 39 disciplines, further divided into sub-corpora by discipline (SCD), field (SCF), and section (SCS), allowing researchers to compare language differences across disciplines and across paper segments (such as introduction versus discussion). For postgraduate students writing papers in English, this amounts to a searchable reference system for "academic English usage."
An anti-corruption corpus: what exactly does the HKCCP contain?
Among this stack of finance- and engineering-heavy corpora, the Hong Kong Corpus of Corruption Prevention (HKCCP) is the small specimen that best illustrates the ambition of "professional communication": it turns the linguistic microscope on the corruption-prevention documents of the Independent Commission Against Corruption.
According to the HKCCP's official page※, the corpus contains 49 documents from the ICAC's Corruption Prevention Advisory Service, plus Hong Kong's Prevention of Bribery Ordinance (Cap. 201) and the UK's Bribery Act 2010. The entire corpus comprises 567,599 words and 10,932 distinct word forms; the page shows it was last updated on 12 August 2018.
This combination is telling: it places the two anti-corruption laws of Hong Kong and the UK, together with the ICAC's operational prevention guidelines, in the same searchable repository, so researchers can compare word by word — how terms like "conflict of interest" and "advantage" in the anti-corruption field are used in legislative language versus practical guidelines, what verbs they collocate with, and in what sentence structures they appear. This is the quintessential question of English for Specific Purposes: not "what does this word mean" but "how is this word used in this profession."
The willingness of a department to put effort into building an anti-corruption corpus of under 600,000 words is itself a declaration of values: it affirms that even "the English of corruption-prevention documents" deserves study as an independent professional genre. This follows the same logic as building the 7.2-million-word financial corpus — every professional community has its own variety of English, and every one of them deserves to be authentically recorded and described.
From corpus to classroom: how does the data feed back into professional English teaching?
If a corpus simply sat in a research centre, it would be no more than an academic asset. The institutional design of the Department of English and Communication is to let it flow all the way into the classroom. The department's taught master's programme — the MA in English Studies for the Professions — is precisely the teaching outlet for this "real professional language" tradition. One of its four specialisms is directly named "English for the Professions," aimed at practitioners and institutional administrators in accounting, engineering, law, and other industries. For the overall structure of PolyU's taught postgraduate programmes and the general English language requirement, see Programmes and General Education Overview.
This path from data to podium is institutionally self-consistent: the department first uses the RCPCE's industry corpora to map out "how engineers and financial professionals actually write English," then turns those findings into teaching materials for the next cohort of students entering those industries. According to the department page on PolyU Scholars Hub※, the department's research output has long been anchored in language and professional communication — and it is precisely this main line that supplies a steady stream of content for the curriculum.
Also worth noting is the methodological discipline of this tradition. Its default stance is "corpus-driven": conclusions must grow out of real data, not be deduced from intuition or convention. From the HKCSE's sentence-by-sentence intonation transcription, to the concgram's exhaustive enumeration of collocational permutations, to the word-by-word collection of the six industry corpora, the same discipline runs through everything — first observe what language actually is, then talk about how it should be taught. This resonates strongly with PolyU's overall "applied" institutional character: even when researching language, the point is to study how language does its job in the real world.
This "data–classroom" loop also distinguishes the department's teaching from the "remedial" English courses of a typical language centre. When students learn to write an engineering tender, a financial-industry email, or a compliance document, what lies behind it is a set of genre features inductively derived from real engineering and financial corpora — teaching not "the correct English of the textbook" but "how people in this profession actually write." Vague-language research tells students the social functions of imprecise expressions in intercultural settings; speech-act research lets them see how requests, concessions, and confirmations cycle through a meeting. Every one of these findings read out of the HKCSE and the industry corpora can ultimately be converted into concrete, teachable classroom skills. For a department whose mission statement says "succeed in the workplace," this pipeline from corpus to workplace skills is its very reason for existing.
| Stage | What the department does | Institutional vehicle |
|---|---|---|
| Collection | Collects authentic professional texts and speech from various industries | RCPCE industry corpus family |
| Description | Analyses collocation and genre using concgram and other methods | ConcGram software, SSCI papers |
| Teaching | Converts findings into professional English course content | MA in English Studies for the Professions (four specialisms) |
| Support | Provides English support for non-humanities students university-wide | General English language requirement |
What does this corpus tradition mean in PolyU's humanities landscape?
Placed within the overall structure of PolyU's Faculty of Humanities, the Department of English and Communication and its sister units — Chinese and Bilingual Studies, Language Sciences and Technology, and others — are not each speaking their own language. They share a hidden through-line of "language technology and applied linguistics." Corpus, discourse analysis, language data — these methods flow between departments, constituting the technical foundation of the Faculty of Humanities' overall rise in linguistics in recent years (for the faculty overview, recent rankings, and appointments, see Faculty of Humanities deep-dive dossier).
The Department of English and Communication's distinctive contribution to this through-line is doing "professional/occupational English" thoroughly and deeply. Other units may lean toward Chinese linguistics, bilingual translation, or speech therapy; this department holds the ground of "how English is used in the workplace," and has materialised that ground with an entire corpus family. Its value lies not in the number of researchers or the size of its budgets, but in the clarity of its methodological identity: it knows it is a "corpus-driven, professionally-oriented communication" outfit.
This also explains why the department deserves its own dedicated file. In an applied university known for engineering, design, hotel management, and health sciences, an English department can avoid marginalisation and find an irreplaceable position precisely because it welds together a humanities method (language description) with an applied purpose (workplace communication). It proves that at PolyU, even the study of English can be a thoroughly applied discipline — one that speaks through real data and finds its ultimate purpose in the workplace.
Putting the tradition on a timeline: from "Department of English" to "Department of English and Communication"
Arranging the above milestones by year reveals a clear trajectory of "corpus first, identity later." The corpus was not a prestigious addition acquired after the department became famous; it is the process by which the department, step by step, defined what it is. The word "Communication" in the department's current name is precisely the self-understanding that grew, piece by piece, out of real professional language over these three decades.
| Date | Milestone | Institutional significance |
|---|---|---|
| Mid-1990s – early 2000s | HKCSE compiled (then under "Department of English") | Establishes the corpus-driven method and the intercultural corpus foundation |
| 2008 | HKCSE prosodic monograph published (SCL 32) | First application of Brazil's Discourse Intonation to a real spoken corpus |
| 2009 | Concgram paper (Applied Linguistics) and ConcGram 1.0 software | Self-coined method reaches an SSCI top journal and commercial release |
| 2010s | RCPCE industry corpus family expands (finance, engineering, surveying, research articles) | Method spread across specific industries, reaching tens of millions of words |
| 2018 | Hong Kong Corpus of Corruption Prevention (HKCCP) updated | Professional genre research refined down to corruption-prevention documents |
| Present | Department named "English and Communication" | "Professional communication" identity institutionalised and fixed |
Note on figures: The milestones above are compiled from the publication catalogues of each corpus, software release dates, and update dates shown on RCPCE pages. The exact year of the department's renaming is not individually documented on the official pages; here, "Department of English" is shown as the historical name during the corpus-compilation period, and "Department of English and Communication" as the current name, presented side by side.
This timeline also answers a question that is easy to overlook: how PolyU's English team managed to leap up the QS linguistics rankings (see Faculty of Humanities deep-dive dossier for the ranking and recent appointments). The answer does not lie in any single appointment or a lucky year, but in three decades of steadily accumulated corpus infrastructure and methodological reputation. A jump in the rankings is the belated return on long-term investment — not a castle in the air.
Frequently asked questions
Can only scholars access the department's corpora?
No. According to RCPCE's industry corpus portal page※, the Hong Kong Financial Services Corpus, the Hong Kong Engineering Corpus, the Corpus of Research Articles, and the Hong Kong Corpus of Corruption Prevention are all open to the public via online search interfaces, offering standard search, part-of-speech search, and advanced search, among other query modes. They are positioned as public resources for teaching and research, usable by professionals, researchers, teachers, and students alike — not closed internal databases.
What is the difference between the HKCSE and the later industry corpora?
The Hong Kong Corpus of Spoken English (HKCSE) is a general spoken corpus of approximately 1 million words, with four sub-corpora of equal size and rare prosodic/intonational transcription (per the SCL 32 catalogue※). The RCPCE's later corpora, by contrast, are a set of industry-specific corpora — such as finance (approx. 7.34 million words) and engineering (approx. 9.2 million words) — consisting mainly of written professional texts, segmented by industry. The former established the method and the reputation; the latter spread the method across specific industries.
Is this department centred on literature, translation, or something else?
None of these, in the traditional sense. According to the department's official overview page※, the Department of English and Communication is positioned as English education centred on workplace skills, with its main research line being "Language and Professional Communication" — that is, applied linguistics, discourse and genre analysis, and corpus research in professional contexts. Literature has been compressed into an optional specialism in the taught master's programme, while translation is mainly accessed through cross-departmental double majors (such as the English and Applied Linguistics and Linguistics and Translation double major); the main translation and interpreting units sit elsewhere in the Faculty of Humanities. In other words, the department's distinctiveness lies not in literature or translation, but in the industry corpora and professional-communication methodology it has built up over three decades.
Is the concgram a PolyU original?
According to John Benjamins' software catalogue※ and the IJCL paper※, the concgram term, method, and its eponymous search engine (ConcGram 1.0, published 2009) were all proposed and developed by the PolyU team (Cheng, Greaves, Warren, and colleagues). The systematic treatment was published in Applied Linguistics (vol. 30, issue 2) and funded by Hong Kong's Research Grants Council (PolyU 5459/08H). It is fair to say this is one of the department's most emblematic methodological originals.
Sources
- Department of English and Communication official overview page※ — department mission, positioning, five research areas.
- RCPCE industry corpus portal page※ — overview of the profession-specific corpus family and public search.
- Hong Kong Financial Services Corpus (HKFSC) page※ — financial corpus size (7,341,937 words), 25 text types.
- Hong Kong Corpus of Corruption Prevention (HKCCP) page※ — anti-corruption corpus contents, size (567,599 words), and update date.
- Cheng, Greaves & Warren, A Corpus-driven Study of Discourse Intonation (SCL 32, 2008)※ — HKCSE compilation, prosodic transcription, Brazil's four Discourse Intonation systems.
- Greaves, ConcGram 1.0: A phraseological search engine (2009)※ — the concgram search engine.
- Cheng, Greaves & Warren, "From n-gram to skipgram to concgram" (IJCL 11:4)※ — conceptual definition of the concgram.
- PolyU Scholars Hub — Department of English and Communication※ — main research output lines.
- Cross-reference: Module 01 · Faculty of Humanities deep-dive dossier (including current overview of the Department of English and Communication, QS rankings, and the McEnery appointment).
This dossier is the single-department institutional study within Module 01 (focus on the corpus tradition), complementary to the Faculty of Humanities overview. Corpus sizes and text types follow the current records of the RCPCE and publication catalogues; word counts are noted with their basis and time point.
Sources · verify independently
- OfficialPolyU Department of English and Communication — Department Overview
- OfficialRCPCE Profession-specific Corpora(行业语料库入口)
- OfficialHong Kong Financial Services Corpus(HKFSC)
- OfficialThe Hong Kong Corpus of Corruption Prevention(HKCCP)
- AcademicCheng, Greaves & Warren, A Corpus-driven Study of Discourse Intonation(SCL 32, John Benjamins 2008)
- AcademicGreaves, ConcGram 1.0: A phraseological search engine(John Benjamins 2009)
- AcademicCheng, Greaves & Warren, From n-gram to skipgram to concgram(IJCL 11:4)
- OfficialPolyU Scholars Hub — Department of English and Communication