Department of English and Communication – In-depth File: Professional Communication, Translation, and Corpus Linguistics
Module: 01 Academic · Sub-file: An Institutional Investigation of the Department of English and Communication (Corpus Tradition Special) It is not an everyday occurrence inside a Chinese university for an English department to bypass Shakespeare and instead record authentic English from Hong Kong bankers in meetings, engineers drafting tenders, and the ICAC writing corruption-prevention guidelines. Without rehashing the department’s current overview, QS rankings, or personnel appointments (cf. Faculty of Humanities In-depth File (incl. ENGL special section)), this article sets out from institutional evidence to answer an older question: what made this department choose corpus linguistics as its foundation. Materials are drawn principally from PolyU’s official pages, the Research Centre for Professional Communication in English (RCPCE) corpus pages, and academic publisher catalogues. For the department’s place within the University’s overall faculty and department structure, see Faculty and Department Overview.
The one-sentence conclusion: The Department of English and Communication at The Hong Kong Polytechnic University stakes its identity on “language excellence in professional contexts”; since the HKCSE in the 1990s it has built a family of industry-specific corpora spanning finance (~7.34 million words), engineering, and anti-corruption, and it originated the concgram methodology.
Why “Department of English and Communication” Rather Than “Department of English”?
The name of The Hong Kong Polytechnic University’s Department of English and Communication (ENGL) already tells you where it parts company with a conventional English department. Most university English departments take Anglo-American literature as their core, with courses revolving around poetry, fiction, and literary criticism. PolyU’s unit, by contrast, anchors itself in Communication—how language is actually used in real professional settings, not how it is beautified in the literary canon.
This orientation is written into the department’s official mission. According to the department overview page※, the department’s primary objective is to provide “first class English language education with a focus on the skills and expertise needed to succeed in the global and the local workplace.” A single mission statement underlines “workplace” twice, nailing its academic positioning firmly to the applied side.
In institutional terms, the orientation is not mere rhetoric but a built structure. The department organises research into five areas, the foremost of which is “Language and Professional Communication,” covering lexicogrammar, discourse, genre, and pragmatics in professional contexts, as well as organisational communication, intercultural workplace communication, and English for Specific Purposes (ESP). Literature and language arts have not vanished, but they have been slimmed down to an optional specialism within a taught master’s programme rather than remaining the department’s centre of gravity.
To grasp where this “workplace-based” identity comes from, one must return to the 1990s—to a corpus project that set out to capture the English actually spoken by Hong Kong people exactly as it was. That corpus is not just a dataset; it is the methodological starting point of this department’s identity: do not infer what language ought to be from a textbook; observe what it actually is from authentic data.
It All Began with a Corpus That “Listens”: How Was the HKCSE Built?
The Hong Kong Corpus of Spoken English (HKCSE) is the foundational project of this tradition. According to academic publisher John Benjamins’s catalogue entry※, the corpus was compiled from the mid‑1990s to the early 2000s by three scholars—Winnie Cheng, Chris Greaves, and Martin Warren—inside what was then simply the “Department of English” at the University. The name “English and Communication” had not yet been adopted; the corpus was born before the name.
The HKCSE comprises roughly one million words, divided into four equal sub-corpora: academic, conversation, business, and public discourse. All recorded speakers were adults, typically native speakers of Cantonese or English—a design choice that deliberately preserves Hong Kong’s intercultural context. What was captured was not “Standard British English” but the English that Hong Kong people actually speak in real situations. The corpus therefore records a site of interlanguage and intercultural communication, not a model of normative English.
The most unusual feature of the HKCSE is that it underwent both orthographic transcription and prosodic transcription—not only were the words written down, but so was how they were said: where the prominence fell, the rise and fall of intonation, the pitch movements. According to the same catalogue entry, the 2008 monograph A Corpus-driven Study of Discourse Intonation (Studies in Corpus Linguistics, vol. 32) was the “first” work to systematically apply linguist David Brazil’s four systems of Discourse Intonation—prominence, tone, key, and termination—to a corpus of authentic, naturally occurring spoken English.
This meant the PolyU team had to handle not quiet written texts but noisy, overlapping, accented, emotionally charged real speech. They even had to build their own tools just to search this intonation‑annotated material. It was precisely this path of “first acquire real data, then let the data force new methods” that propelled the department towards its next original contribution: the concgram.
What Is a Concgram? How Did PolyU Scholars Redefine “Collocation”?
Corpus linguistics had long been counting n‑grams (contiguous sequences of n words, such as “in order to”). Yet in real language, a set of words that frequently co‑occur often neither adjoin one another nor appear in a fixed order—for instance, “play” and “role” can be separated by “an important,” “a key,” “a decisive,” and can even be flipped into “the role played by.” Conventional n‑grams fail to capture this loose but genuine co‑occurrence.
The solution proposed by the PolyU team is called the concgram. As defined in From n-gram to skipgram to concgram※ (published in the International Journal of Corpus Linguistics, Volume 11, Issue 4), a concgram is a co‑occurrence unit generated by the association of two or more words that accounts for all permutations of constituency variation and positional variation. In plain terms, it treats every instance where the same words recur together—regardless of how many words come between them or which appears first—as belonging to a single collocational unit.
The concept was subsequently embedded in software. According to John Benjamins’s software catalogue entry※, Chris Greaves’s “ConcGram 1.0: A phraseological search engine,” officially released in 2009, became the first title in the “Studies in Corpus Linguistics: Software” series. The engine can automatically detect co‑occurrences exhibiting both constituency and positional variation, making it, in the publisher’s words, “uniquely suited to revealing the full phraseological profile of a text or corpus.”
The methodology’s weight also carries endorsement from a top‑tier journal. Cheng, Greaves, Sinclair, and Warren’s Uncovering the extent of the phraseological tendency: towards a systematic analysis of concgrams appeared in 2009 in Applied Linguistics (Volume 30, Issue 2); the related research was supported by a Hong Kong Research Grants Council (RGC) grant (project no. PolyU 5459/08H). Taking a self‑coined term into a leading SSCI journal, turning it into a commercially released piece of software, and then winning competitive research funding—this is a complete chain that “institutionalises” a method, and every link of that chain falls inside this department.
It is worth spelling out what assumptions this method challenges. The traditional view of language leans towards treating sentences as “syntactic skeletons into which vocabulary is slotted”; corpus linguistics has discovered that language relies heavily on strings of semi‑fixed chunks—what the late John Sinclair, a founder of the field, called the “idiom principle.” The phrase “phraseological tendency” in the paper’s title is exactly a continuation of this thesis: language is far more “pre‑packaged” than grammarians imagine. The concgram’s contribution is that it provides this abstract thesis with a tool that can exhaustively enumerate and quantify that pre‑packaging—it turns “how formulaic is language?” from a philosophical assertion into a measurable corpus fact. Sinclair’s name appearing among the paper’s authors also plugged the PolyU team straight into the orthodox lineage of international corpus linguistics.
This method can also be applied directly to industry‑specific corpora. According to a publicly available RCPCE analytical report (Mike Report※), researchers once used concgrams to examine the collocational network of terms such as “attributable” and “shareholders” inside the Hong Kong Financial Services Corpus—aiming a home‑grown method squarely at real financial texts to see exactly how words cluster in professional English. Here method and data converge: a self‑built tool applied to a self‑built industry dataset, answering the real questions of “professional communication.”
Once the HKCSE Was Recorded, What Did PolyU Scholars Read Out of It?
The true value of a corpus lies in how much research can grow out of it. After it was completed, the Hong Kong Corpus of Spoken English (HKCSE) became a rich seam that the department mined repeatedly over two decades—it proved that “recording real language” is not the destination, but the starting point for a chain of discoveries.
The first seam was vague language. In “The Use of Vague Language Across Spoken Genres in an Intercultural Hong Kong Corpus” (included in an academic edited volume, Springer chapter page※), PolyU scholars took representative samples from the four HKCSE sub-corpora—academic, business, conversation, public—and analysed how Hong Kong speakers in intercultural settings use expressions like “and so on,” “or something,” “about”—words that appear imprecise yet carry real social functions. Vagueness is not a language flaw but a strategy for lubricating intercultural communication—a finding that could only be made with an authentic corpus. This line of research can be traced even earlier, to “The use of vague language in intercultural conversations in Hong Kong,” published in English World-Wide (Volume 22, Issue 1).
The second seam was the pragmatic function of a single word. Along the line of work represented by research in the International Journal of Corpus Linguistics (Volume 6, Issue 2)※, the PolyU team once tracked the use of the word “actually” in intercultural conversation data—how one unremarkable adverb is deployed to signal contrast, correction, or to soften disagreement. Turning one small word into a full research paper is a classic demonstration of corpus linguistics’s ability to find “big truths in small details.”
The third seam pushed the HKCSE towards speech‑act annotation. According to the RCPCE’s HKCSE Speech Act Corpus (HKCSE-SAC) page※, researchers annotated a searchable speech‑act sub‑corpus on top of the HKCSE, identifying 45,372 spoken English discourse segments and 69 speech‑act types. The sub‑corpus was compiled by Dr Seto Wood‑hung Andy as part of his doctoral dissertation. The quantitative findings showed that although meetings, telephone calls, informal office talk, service encounters, Q&A sessions, and interviews represent six distinct interactional genres, the number and variety of unique speech acts, as well as their frequencies and co‑occurrences, display cross‑genre similarities—suggesting that authentic professional spoken language is structurally more consistent beneath the surface than it first appears.
Taken together these three seams demonstrate one thing: the Department of English and Communication did not build a corpus only to leave it on the shelf; around it grew an entire research community spanning vague language, pragmatics, and speech acts. One corpus, twenty years of excavation, dozens of papers and several monographs—that is what gives the “corpus tradition” its genuine weight, and it is the core asset that sets this department apart from an ordinary English teaching unit.
How Large Is the RCPCE’s “Family of Industry Corpora”?
The foundation laid by the HKCSE and concgrams eventually grew into a dedicated research unit: the Research Centre for Professional Communication in English (RCPCE). No longer content with a single general spoken corpus, the Centre set out to collect authentic professional texts from a range of Hong Kong industries—turning “language in professional contexts” from a slogan into a searchable database cluster.
According to the RCPCE’s industry corpora portal※, the Centre gathers authentic texts, discourse, and genres from various Hong Kong professional communities and contexts, building a series of profession‑specific corpora that are publicly accessible and searchable online. What these corpora have in common is that the material comes from the English genuinely produced at work by engineers, financial practitioners, and anti‑corruption officials—not from textbooks or news media. Below, the main members of this “corpus family” are listed by attested size:
| Corpus | Content | Size (metric) | Source Page |
|---|---|---|---|
| Hong Kong Corpus of Spoken English (HKCSE, prosodic version) | Authentic spoken English in academic/conversation/business/public domains, with intonation transcription | ~1 million words | SCL 32 catalogue※ |
| Hong Kong Financial Services Corpus (HKFSC) | 25 text types from the financial services industry | 7,341,937 words | HKFSC page※ |
| Hong Kong Engineering Corpus (HKEC) | Authentic English texts from the engineering profession | ~9.2 million words | RCPCE portal※ |
| Surveying and Construction Engineering Corpus | Texts from surveying and construction engineering | ~5.7 million words | RCPCE portal※ |
| Corpus of Research Articles (CRA) | Research articles across 39 disciplines, with discipline/field/section sub‑corpora | Multi‑disciplinary (see below) | RCPCE portal※ |
| Hong Kong Corpus of Corruption Prevention (HKCCP) | ICAC corruption‑prevention guidelines and related legislation | 567,599 words | HKCCP page※ |
Metric note: The word counts given above are the compiled sizes as recorded on each corpus’s official page and in published catalogues, calculated as cumulative tokens at the time the corpus was built; word counts for the engineering and surveying corpora are given to the nearest integer level (~9.2 million, ~5.7 million) as described in public RCPCE materials; readers should consult the centre’s pages for the most up‑to‑date exact figures.
Even a glance at the numbers reveals the emphasis: finance (~7.34 million words) and engineering (~9.2 million words) are the two thickest slices, corresponding precisely to Hong Kong’s twin pillar industries as an international financial centre and an infrastructure‑intensive city. The distribution of the corpus family itself reads as a profile of the professional English landscape in Hong Kong—which industries the department chose to record reflects its judgement about what kind of language the city runs on.
The Corpus of Research Articles (CRA), meanwhile, extends the reach into academic writing itself. According to RCPCE public materials, the CRA includes research articles spanning 39 disciplines, split into multiple sub‑corpora by discipline (SCD), subject field (SCF), and section (SCS), enabling researchers to compare linguistic differences across disciplines and across different parts of a paper (e.g., introduction versus discussion). For postgraduates writing in English, this amounts to a searchable reference system for the “usage conventions of academic English.”
An Anti‑Corruption Corpus: What Exactly Does the HKCCP Contain?
Among all these corpora weighted towards finance and engineering, the Hong Kong Corpus of Corruption Prevention (HKCCP) is the small specimen that best illustrates the ambition of “professional communication”—it trains a linguistic microscope on the corruption‑prevention documents of the Independent Commission Against Corruption.
According to the HKCCP’s official page※, the corpus contains 49 documents from the Corruption Prevention Advisory Service of Hong Kong’s ICAC, together with Hong Kong’s Prevention of Bribery Ordinance (Cap. 201) and the UK Bribery Act 2010. The full corpus totals 567,599 words and 10,932 distinct word forms; the page shows it was last updated on 12 August 2018.
This composition is telling: it places the anti‑bribery statutes of Hong Kong and the United Kingdom alongside the ICAC’s practical corruption‑prevention guidelines inside a single searchable database. Researchers can therefore conduct word‑by‑word comparisons—how terms such as “conflict of interest” or “advantage” are used in legislative language versus in operational guidance, what verbs they collocate with, what sentence patterns they appear in. This is the archetypal question of English for Specific Purposes research: not “what does this word mean,” but “how is this word used in this profession.”
That a department would invest the effort to build an anti‑corruption corpus of fewer than 600,000 words is in itself a declaration of scholarly values: it asserts that even “the English of corruption‑prevention documents” deserves study as a distinct professional genre. The logic is the same as the one behind building a 7‑million‑word financial corpus—every professional community has its own variety of English, and every variety deserves to be authentically recorded and described.
From Corpus to Lectern: How Does This Data Feed Back into Professional English Teaching?
If corpora only lay inside a research centre, they would remain merely an academic asset. The Department of English and Communication’s institutional design is to channel them all the way into the classroom. The department’s taught master’s degree—the MA in English Studies for the Professions—is the pedagogical outlet for this “real professional language” tradition, and among its four specialisms one is directly titled “English for the Professions,” aimed at practitioners in accounting, engineering, law, and institutional administrators. For the overall architecture of PolyU’s taught master’s programmes and the general English language requirements, see Programmes and General Education Overview.
This path from data to lectern is institutionally self‑consistent: the department first uses the RCPCE’s industry corpora to describe “how engineers/financial practitioners actually write in English,” then turns those findings into teaching materials for the next cohort of students bound for those industries. According to the departmental page on PolyU Scholars Hub※, the department’s research output has long taken language and professional communication as its main line, and it is precisely this line that supplies the curriculum with a steady stream of content.
What deserves highlighting is the methodological discipline of this tradition. Its default posture is “corpus‑driven”: conclusions must grow out of authentic data, not be deduced from intuition or prescription. From the utterance‑by‑utterance intonation transcription of the HKCSE, to the exhaustive enumeration of all collocational permutations by concgrams, to the word‑by‑word collection across six major industry corpora, the same discipline runs through the whole enterprise—first observe what language actually looks like, then decide how it should be taught. This is highly consonant with PolyU’s university‑wide “application‑oriented” ethos: even when studying language, study how language does its job in the real world.
This “data‑to‑classroom” loop also distinguishes the department’s teaching from the “remedial” English classes of a typical language centre. When students learn how to draft an engineering tender, a financial‑sector email, or a compliance document, what underpins the lesson are the genre features induced from real engineering and financial corpora—what is taught is not “correct textbook English” but “how people in this profession actually write.” Vague‑language research shows students the social functions of fuzzy expressions in intercultural settings; speech‑act research lets them see how requests, concessions, and confirmations shuttle through a meeting. These findings, extracted from the HKCSE and the industry corpora, are ultimately converted into teachable, concrete skills. For a department that has written “succeed in the workplace” into its mission, this pathway from corpus data to workplace skill is its very reason for being.
| Stage | What the department does | Institutional vehicle |
|---|---|---|
| Collection | Gathers authentic professional texts and spoken data from various industries | RCPCE industry corpus family |
| Description | Analyses collocation and genre using concgrams and other methods | ConcGram software, SSCI journal papers |
| Teaching | Converts findings into professional English course content | MA in English Studies for the Professions (four specialisms) |
| Support | Provides English language support for non‑humanities students across the whole university | General English language requirements |
What Does This Corpus Tradition Mean Within PolyU’s Humanities Landscape?
Placed back into the overall structure of PolyU’s Faculty of Humanities, the Department of English and Communication and its sister units in Chinese/Bilingual Studies and Linguistic Sciences and Technology are not speaking unrelated languages. Rather, they share a common thread of “language technology and applied linguistics.” Corpus, discourse analysis, and language data—these methods flow among the departments and form the technical chassis behind the Faculty’s recent rise in the linguistics discipline (for the Faculty overview and recent rankings and personnel, see Faculty of Humanities In‑depth File).
The unique contribution of the Department of English and Communication to this shared thread is in doing “professional/occupational English” thoroughly and deeply. Other units may lean more towards Chinese linguistics, bilingual translation, or speech therapy, while this department guards the terrain of “how English is used in the workplace”—and has materialised that terrain with an entire family of corpora. Its value lies not in a large number of researchers or vast funding, but in a clarity of methodological identity: it clearly understands itself as a “corpus‑driven, professionally‑oriented communication” team.
This also explains why the department deserves its own archival file. In an applied university best known for engineering, design, hospitality, and health sciences, an English department has managed not to be marginalised but instead to occupy an irreplaceable position—precisely by fusing a humanities method (linguistic description) with an applied purpose (workplace communication). It has proved that at PolyU, even the study of English can be a thoroughgoing applied discipline—an applied discipline that speaks with authentic corpora and takes the workplace as its destination.
Placing This Tradition on a Timeline: From “Department of English” to “Department of English and Communication”
If the milestones above are laid out chronologically, a trajectory of “corpus first, identity later” becomes clearly visible. The corpus is not a showpiece added after the department became famous; it is the process through which the department, step by step, defined what it is. The word “Communication” in the department’s current name is a self‑understanding that grew, over thirty years, out of real professional language.
| Time point | Milestone | Institutional significance |
|---|---|---|
| Mid‑1990s – early 2000s | Compilation of the HKCSE (within the then “Department of English”) | Laid the corpus‑driven method and intercultural corpus foundation |
| 2008 | Publication of the prosodic HKCSE monograph (SCL 32) | First systematic application of Brazil’s Discourse Intonation to an authentic spoken corpus |
| 2009 | Concgram paper (Applied Linguistics) and ConcGram 1.0 software | Self‑originated method enters a top SSCI journal and achieves commercial release |
| 2010s | Expansion of the RCPCE industry corpus family (finance, engineering, surveying, research articles) | Method scaled out to specific industrial sectors, reaching a total of tens of millions of words |
| 2018 | Hong Kong Corpus of Corruption Prevention (HKCCP) updated | Professional genre research refined to the level of corruption‑prevention documents |
| Present | Department permanently named “Department of English and Communication” | “Professional Communication” identity institutionalised |
Metric note: The milestones above are compiled from corpus publication catalogues, software release years, and the update timestamps shown on RCPCE pages; the precise year of the departmental name change is not listed on official pages, so “Department of English” is used as the historical name during the corpus‑compilation period and “Department of English and Communication” as the current name.
This timeline also answers an easily overlooked question: why this particular English team at PolyU has been able to leap forward in the QS World University Rankings by subject in Linguistics (for those rankings and recent personnel changes, see Faculty of Humanities In‑depth File). The answer lies not in any single appointment or one lucky year, but in the corpus infrastructure and methodological reputation steadily accumulated over three decades—the jump in rankings is the belated return on a long‑term investment, not a castle in the air.
Frequently Asked Questions
Are the department’s corpora accessible only to academics?
No. According to the RCPCE’s industry corpora portal※, the Hong Kong Financial Services Corpus, Hong Kong Engineering Corpus, Corpus of Research Articles, and Hong Kong Corpus of Corruption Prevention are all publicly accessible through online search interfaces offering standard, part‑of‑speech, and advanced query options. They are positioned as public teaching and research resources, available to professionals, researchers, teachers, and students, rather than being closed, internal databases.
What is the difference between the HKCSE and the subsequent industry corpora?
The Hong Kong Corpus of Spoken English (HKCSE) is a general spoken corpus of approximately one million words, with four equal sub‑corpora and the rare addition of prosodic intonation transcription (per SCL 32 catalogue※). The later RCPCE corpora, by contrast, are a set of industry‑specific corpora—such as finance (~7.34 million words) and engineering (~9.2 million words)—mainly composed of written professional texts and partitioned by industry. The former laid the methodological foundation and the department’s reputation; the latter scaled the method out to concrete industrial sectors.
Is the department’s centre of gravity literature, translation, or something else?
It is none of these in the traditional sense. According to the department overview page※, the Department of English and Communication is positioned as English education centred on workplace skills, with its main research line being “Language and Professional Communication”—that is, applied linguistics, discourse and genre analysis, and corpus research in professional contexts. Literature has been slimmed down to an optional specialism within a taught master’s degree. Translation is more commonly accessed through cross‑departmental double majors (e.g., English and Applied Linguistics with Linguistics and Translation), and the primary unit responsible for translation and interpreting lies elsewhere in the Faculty of Humanities. In short, the department’s distinctiveness is neither literature nor translation, but the industry corpora and professional‑communication methodology it has built over thirty years.
Is the concgram a PolyU original?
Yes. According to John Benjamins’s software catalogue entry※ and the IJCL paper※, the term, method, and the eponymous search engine (ConcGram 1.0, released in 2009) were all proposed and developed by the PolyU team (Cheng, Greaves, Warren, et al.). The systematic exposition was published in Applied Linguistics (Volume 30, Issue 2) and supported by an RGC grant (PolyU 5459/08H). One can reasonably call it one of the department’s most emblematic original methodological contributions.
Sources
- Department of English and Communication overview page※ — departmental mission, positioning, five research areas.
- RCPCE industry corpora portal※ — overview of the profession‑specific corpus family and public search access.
- Hong Kong Financial Services Corpus (HKFSC) page※ — financial corpus size (7,341,937 words), 25 text types.
- Hong Kong Corpus of Corruption Prevention (HKCCP) page※ — anti‑corruption corpus contents, size (567,599 words), and last update.
- Cheng, Greaves & Warren, A Corpus‑driven Study of Discourse Intonation (SCL 32, 2008)※ — HKCSE construction, prosodic transcription, Brazil’s four systems of Discourse Intonation.
- Greaves, ConcGram 1.0: A phraseological search engine (2009)※ — the concgram search engine.
- Cheng, Greaves & Warren, From n‑gram to skipgram to concgram (IJCL 11:4)※ — the definition of the concgram concept.
- PolyU Scholars Hub — Department of English and Communication※ — main line of research output.
- Cross‑reference: 01 Academic · Faculty of Humanities In‑depth File (incl. current ENGL overview, QS ranking, McEnery appointment).
This file is a single‑department institutional investigation (corpus tradition special) within the 01 Academic module, complementing the Faculty of Humanities overview. Corpus sizes and text types are based on the current records of the RCPCE and published catalogues; word counts are noted with their metric and time point.
Sources · verify independently
- OfficialPolyU Department of English and Communication — Department Overview
- OfficialRCPCE Profession-specific Corpora(行业语料库入口)
- OfficialHong Kong Financial Services Corpus(HKFSC)
- OfficialThe Hong Kong Corpus of Corruption Prevention(HKCCP)
- AcademicCheng, Greaves & Warren, A Corpus-driven Study of Discourse Intonation(SCL 32, John Benjamins 2008)
- AcademicGreaves, ConcGram 1.0: A phraseological search engine(John Benjamins 2009)
- AcademicCheng, Greaves & Warren, From n-gram to skipgram to concgram(IJCL 11:4)
- OfficialPolyU Scholars Hub — Department of English and Communication