Redbird Labs

Concept-Conditional Cross-Tradition Binding

An open natural-language-processing method for a challenging question: when unconnected traditions seem to describe reality the same way, are they truly converging, or does it only look that way in translation?

What is CCB?

Concept-Conditional Cross-Tradition Binding (CCB) is a natural-language-processing method for testing convergence claims, the idea that two unconnected bodies of text are really “saying the same thing.” Instead of taking that on faith, CCB turns it into a measurable statistic in semantic embedding space.

The method converts passages into high-dimensional vectors and asks a precise question: once you condition on a shared structural concept, are passages from unrelated traditions more similar to one another than chance would predict? Crucially, it is bias-aware, built from the ground up to separate genuine structural convergence from artifacts of shared vocabulary or a common translation language.

It's Redbird's first public research project, released open source with the paper, corpora, and complete analysis pipeline.

How it works

Embeddings, not keywords

Passages become vectors in a high-dimensional semantic space using modern embedding models. The method compares where ideas sit in relation to one another, not whether texts happen to share the same words.

Conditioned on concepts

Rather than assuming traditions share vocabulary, CCB fixes a set of structural concepts and measures whether unrelated traditions position comparable ideas in structurally similar relationships.

Bias-aware controls

Every measured effect is checked against baselines and null models, so apparent convergence has to survive controls before it counts as signal rather than noise.

Cross-lingual validation

The same tests are run across languages to check whether a pattern is a real structural feature or just an artifact of English translation acting as a common confounder.

An application to mysticism

The first proving ground for CCB is a debate that has run for over a century: do the world's contemplative and mystical traditions converge on a shared description of reality, or does each construct its own? It's an ideal stress test: the traditions are largely independent, the texts span millennia and many languages, and the convergence question has been argued philosophically for decades without a common measuring stick.

The project treats this literature as a multi-millennium dataset of hypotheses about reality and follows a phased, pre-registered design. Four exploratory phases have progressively tightened the controls, moving from a curated corpus to whole books to multiple translators to original-language texts in nine languages. Each one removes a way the apparent convergence could be an artifact rather than a real signal.

Completed exploratory phases

Phase 0: Curated corpus

An exploratory first pass over 143 passages spanning 23 traditions, used to develop and sanity-check the method.

Phase 1a: Whole books

920 chunks (~2.85M tokens) from 20 complete books across 11 traditions, testing the binding concepts on full, non-paraphrased text.

Phase 1b: Multiple translators

Re-runs on the Bhagavad Gita and Tao Te Ching with three translators each, to bound how much a single translator's voice drives the result.

Phase 2: Original languages

Original-language texts across nine languages (including Chinese, Arabic, Hindi, Hebrew, Greek, and Japanese), removing English translation as a shared confounder.

Next · pre-registered

Phase 3a: Confirmatory test

The project's first confirmatory step. It isolates genealogical contact (shared lineage between traditions) as the last major confound, comparing genealogically independent Axial-Age spheres (pre-Buddhist China and pre-contact Greece) in their original languages. The latest release formally pre-registers its hypotheses; the experiment is planned, not yet run.

Why it matters for NLP

A reusable convergence test

CCB isn't limited to one domain. It's a general framework for testing whether any two bodies of text converge on shared structure, useful anywhere “these are basically the same” needs to be checked rather than assumed.

Translation artifact vs. real structure

By validating across languages, the method helps separate patterns that are genuine features of the ideas from patterns that are really just properties of the translation.

Fuzzy claims made testable

It operationalizes a philosophical hypothesis as a concrete, falsifiable computational claim, turning an argument that resisted measurement into something you can actually run.

Rigor built in

Pre-registration, baselines, null models, and multiple embedding backends make the results defensible, not just suggestive.

Open by design

We consider this project open scientific research, so CCB is meant to be shared. The paper, corpora, analysis pipeline, and result tables are all public, and the code is MIT licensed, so anyone can inspect the method, reproduce the results, or adapt it to their own domain.

We think that's how methodology should work: contribute a tool the field can build on, and let the results stand up to scrutiny in the open.

Need rigorous, measurable AI on your most challenging question?

The same discipline behind CCB (turning fuzzy claims into testable statistics, controlling for bias, and validating every result) is what Redbird brings to client work in NLP and custom AI.

Get in touch