Explore

From access to analysis

Despite its capacity to enrich African history and the study of Muslim societies through alternative modes of knowledge production, the Global South — including Africa — remains under-represented in the Digital Humanities (DH). The Islam West Africa Collection (IWAC) pushes beyond preservation and access to open new lines of enquiry on Islam and Muslim life in West Africa. Rather than serving digitised sources without interpretation (Robertson & Mullen 2021), the pages below analyse the IWAC dataset with computational methods, taking an exploratory approach to media representations of Islam and Muslims and to the intellectual and translocal histories that shape them.

Distant reading

"Distant reading", introduced to literary studies by Franco Moretti, contrasts with the intensive focus on individual texts of "close reading". It uses computational techniques to analyse large textual datasets, detecting patterns and themes that traditional qualitative methods may overlook. With more than 28 million words — over 300 books' worth of text — IWAC has long passed the point where any single reader could survey it whole.

The method has its limitations. Algorithmic modelling can obscure context, tone and meaning that close reading recovers, and the quality of optical character recognition affects everything built on top of it. Distant reading is therefore best combined with close reading and domain expertise. These pages take that mixed-methods approach: every visualisation explains its method in plain language, values produced by AI models are explicitly labelled as such, and nearly every chart links back to the underlying documents so that a pattern spotted at a distance can be verified up close.

How to use these pages

Each page is an analytical instrument: filter, sort, switch views, hover for context, and click through to the sources. Start anywhere — but if you are new to the collection, start with the overview.

Start with the collection

  • Collection overview — the whole archive at a glance: how it has grown, what it contains, which countries, languages and decades it covers, the people, organisations and places it mentions most, and how long each newspaper run extends.

People, places and connections

  • Entity index explorer — the curated index behind the collection: some 4,400 people, organisations, places, subjects and events, with how often and when each appears — and how subject and place tagging has shifted across six decades, including rising and falling subjects and the geography of the press's attention over time.
  • Entity networks — who appears alongside whom: a network of co-occurring people, organisations, events and places across the corpus, and a closer study of how six major Islamic organisations are written about in each other's company.
  • Places — where the collection looks: every geocoded place, sized by attention. Pick any person, organisation, event or subject to see its geography, or focus on a single country.

Themes and language

  • Topics — thirty themes surfaced by topic modelling, their vocabulary, their trajectories across time, countries and newspapers — and a "semantic landscape" that places all 12,000+ articles on a single map, where proximity means textual similarity.
  • Words over time — an n-gram-style viewer tracking any of the corpus's 5,000 most frequent terms year by year, alongside a worked case study: the vocabulary of fear (terrorisme, extrémisme, intégrisme…), annotated with the historical events that drove it.

The press itself

  • Compare newspapers — any two newspapers or country corpora side by side: volume, subjects, vocabulary, geography and tone.
  • Inside the press — the press as an institution: who signed the articles (and the rise of the byline), how the prose read (readability, lexical richness, article length), and how copy travelled between papers, traced through near-duplicate reprints. The reprint pairs are computer-proposed and should be verified, not taken as fact.
  • Press attitudes (AI) — three large language models independently assess each article's tone towards Islam and Muslims. The page shows tonal shifts over time, by country, topic, and newspaper, as well as highlighting areas where the models disagree. Each panel is an AI-generated assessment and is labelled as such.

Beyond the newspapers

  • Islamic publications — the Islamic periodical press on its own terms: 1,500 issues across 25 titles, their runs and holdings, languages and subjects, and a semantic map of issues by their tables of contents.
  • References dashboard — the scholarship around the collection: 860+ academic references, their authors, publishers, subjects and co-authorship networks.

Ask, rather than browse

Ask AI — a conversational interface to the collection. Ask a question in plain language and it will search the corpus and provide answers with linked sources. This is intended to complement the pages above, not replace the need to read them.

Every article, person, organisation, place and periodical issue also carries its own dashboard on its item page, so the instruments above continue at the level of the individual document.

Methods and reuse

All pipelines behind these pages — code and prompts alike — are freely available on GitHub for scrutiny and reuse. For how the collection's metadata is produced, including the AI-assisted enrichment workflow, its risks, and the project's minimal-computing principles, see AI metadata enrichment and Optical character recognition under About.