Artificial intelligence

AI, Culture & Power

Writing on artificial intelligence from a Black cultural and spiritual vantage point — model bias, ancestral data sovereignty, archival restoration, and authorship.

What is Robert Shumake's work on AI?

Robert Shumake writes and teaches about artificial intelligence from a Black cultural and spiritual vantage point. His central argument is that AI systems inherit the archives they are trained on, and that most of those archives underrepresent Black life, African traditions, and community-held knowledge. His work covers four areas: how bias enters models through data and evaluation, how communities can hold sovereignty over cultural and ancestral data, how AI changes publishing and historical restoration, and how writers and practitioners can use AI tools without giving up authorship.

Credential

Harvard Data Science Initiative Agentic AI Intensive

Certificate of Completion from Harvard Data Science Review (verification UTHW-KVJS), June 9–25, 2026. The required capstone rebuilt the curriculum for African entrepreneurs and incarcerated learners, and launched the Smart Money Black AI Executive Program with the National Business League.

Read about the Harvard program and what came out of it

Five working principles

A model is an argument about the archive
Every model encodes a claim about what counts as knowledge. When the training corpus is thin on Black newspapers, oral tradition, and diaspora scholarship, the model is not neutral — it is fluent in one canon and hesitant in another.
Representation without governance is decoration
Adding examples to a dataset does not transfer power. The question is who decides what is collected, who can revoke it, who benefits from the output, and who is accountable when a system is wrong about a person.
Sacred and community knowledge is not public domain
Material can be legally scrapeable and still be culturally restricted. Initiatory, ceremonial, and family knowledge carries conditions of transmission that a crawler cannot read and a license cannot capture.
Restoration is an AI-native task
Digitizing, deskewing, transcribing, and cross-indexing decaying Black newspapers is exactly what machine systems are good at — and it is the fastest way to widen the archive that future models learn from.
Authorship is a discipline, not a setting
Tools accelerate drafting, research, and indexing. Judgment, sourcing, and voice stay with the human who signs the work. The line is kept by process, not by a disclosure sentence at the bottom of a page.

The four areas

Recent scholarship

Shurooms: The First Ethnomycological Study from Within the African Initiatory Tradition

300+ guided psilocybin ceremonies documented by an initiated lineage-holder across three spiritual traditions. Indexed on Zenodo (CERN) with DOI 10.5281/zenodo.21349908.

Read about the study, the Observer Problem, and the three lineages

The books behind this work

7 titles in the catalog deal with artificial intelligence, ancestral intelligence, and the technology economy.

See all AI books and which essay each one pairs with

Common questions

Why is AI biased against Black people?

Because models learn from archives that underrepresent Black publications, dialects, and scholarship while overrepresenting Black people in criminal-justice and risk-scoring text. The imbalance is then hidden by benchmarks that report average accuracy instead of per-group accuracy.

Can AI bias be fixed by adding more diverse data?

More data helps but does not resolve it. Without changes to labeling practice, subgroup evaluation, and who governs the dataset, additional data is absorbed into the same pipeline that discounted the community in the first place.

What is cultural data sovereignty?

It is a community's authority over how knowledge produced by and about it is collected, stored, used for AI training, interpreted, and monetized — including the right to refuse and the right to withdraw.

Can sacred or ancestral knowledge be protected from AI training?

Partially. Copyright covers fixed expression, not tradition, so practical protection comes from governance: keeping restricted material off the public web, publishing explicit training terms, tiering access, and negotiating collectively rather than individually.

How is AI used in historical archive restoration?

For image cleanup, layout analysis, optical character recognition, handwriting recognition, and entity linking across issues. Humans verify low-confidence regions and proper nouns, since machine transcription produces confident errors on degraded type.

Why does digitizing Black newspapers matter for AI?

Undigitized material is not indexed, not cited, and not present in training corpora, so it is missing from the systems people now use to answer historical questions. Digitization directly widens the archive future models learn from.

Should authors use AI to write books?

Use it for mechanical stages — transcription, indexing, cross-referencing, consistency checks — and keep thesis, sourcing, and final prose with the author. The dividing line is whether the tool can detect its own errors, and for factual claims and voice it cannot.

Why do AI tools invent citations?

Because they generate plausible continuations rather than retrieve verified records, and a fabricated citation is as plausible as a real one. The only reliable safeguard is refusing to cite a source the author has not opened.

Related: about Robert Shumake · the book catalog · full FAQ