The space of antibody reactivities modelled by mimotope libraries - IgOme
Apart from the much more popular repertoire sequencing (AIRR-Seq)
What is IgOme?
The term IgOme was coined in a paper from the Jonathan Gershony’s lab1. It consists of bulk mimotope selection from a random peptide phage display library, adsorption on non-specific monoclonals, and NGS of the regions coding for the peptide inserts. This technique can be coupled with bioinformatic analyses of the mimotope libraries for finding of clusters and motifs. It also helps infer properties of the global shape of the reactivity space.
Why this is more than a reframing
Three things follow that the binary picture cannot give:
- Polyreactivity becomes a measurable quantity, not a nuisance — an entropy, not an error bar.
- Mimotope arrays become samples from the space of peptides. A microarray of peptides probes the landscape \(E(\varepsilon)\) at many points at once; the measured reactivities are an empirical sketch of the distribution. Scalable mimotope libraries make this sampling practical at the scale of the public repertoire2.
- The repertoire becomes an ensemble of distributions, opening the door to genuinely statistical-mechanical questions about the population of antibodies as a whole.
None of this is settled. The epitope space is not obviously enumerable, the parameters are not yet estimated or measured, and whether the equilibrium reading is the right one in a dynamic immune system is an open question. But as a way to organise mimotope and microarray data — and as a bridge to the repertoire-physics program — treating specificity as a shape has been more productive than treating it as a switch.
Viewing specificity as a distribution of binding energies may seem hard to reconcile with negative selection. Each monoclonal antibody can select thousands of short peptides from a random peptide library with a biologically relevant affinity. These peptides are found to span the entire peptide space. If each individual reaction can have biological consequences, then the probability of an antibody surviving negative selection would become negligible. Indeed, up to 70% of the early immature B cells in the bone marrow are self-reactive3 and most of them are eliminated by negative selection.
What a panned library is a sample of
The framework needs the density of epitope states \(g(\varepsilon)\), and this page proposes to sketch it from mimotope libraries. It is worth being explicit about what a panning experiment actually samples, because it is not \(g\). It is
\[g(\varepsilon)\times K_{\text{sel}}(\varepsilon)\times F(\text{seq}),\]
truncated below detection: the density of states, times a selection kernel, times a propagation fitness that has nothing to do with binding. The third factor is not a small correction. Sequencing of naive and amplified Ph.D.‑7 libraries found that against a nominal \(1.3\times10^{9}\) NNK\(_7\) diversity only 72% of the naive library consisted of singletons, where Poisson sampling predicts over 99%; the peptide HAIPYRH was present at over 2,000 copies before amplification and over 68,000 after, and it has been reported as a hit against 13 unrelated targets, with LPLTPLP appearing in 11 published screens and SILPYPY in 64. The high‑abundance end of a panned library is therefore contaminated by sequences that propagate well in E. coli, and that is precisely the end that sets \(\varepsilon_{min}\), \(\Delta\), and every participation ratio. Our shape statistics are computed from the most contaminated part of the distribution. Lot‑to‑lot compositional differences compound this5, as do target‑unrelated peptides that bind plastic, blocking agent or capture reagent rather than the antibody6.
Three controls follow, and we should hold ourselves to them. Sequence the naive library and use its profile as an explicit reference measure \(q_0\); the estimand is then never the raw selected frequency but the tilt \[\log[q_n(x)/q_0(x)]\], which cancels library composition bias exactly. This is the control we use widely in our studies. One would Prefer emulsion amplification, which largely preserves abundances when it is available, but it poses another set of technical restrictions.
Two further limits on what can be inferred. Because only exceedances above a detection threshold are observed, the estimable object is a tail, and the right inferential frame is peaks‑over‑threshold rather than block maxima. This is conditional on exceeding a high threshold, Gaussian excesses converge to a generalized Pareto law and the number of exceedances is approximately Poisson. The Gumbel law applies to the single strongest binder, not to the multiset of everything selected — a distinction that matters exactly where we care, in how many near‑degenerate strong binders exist. And the claim that the bulk of \(g\) is Gaussian is not testable from panning data at all, since the bulk is never observed. Panning identifies the tail index of \(g\), not \(g\). What follows is that \(Z_{ep}\), \(S\) and \(N_{eff}\) computed over an observed library are functionals of the tail conditioned on the threshold, and are comparable across antibodies only at matched threshold and matched coverage.
How can these views be reconciled:
- The affinities for small peptides are below the threshold for negative selection, the epitopes above that threshold are less frequent.
- The tolerogenic signals have typically been found to depend on high avidity7 .
- The accessible self-antigens, presented at sufficient concentration and local density, are probably many orders of magnitude fewer than the random peptide species in a phage library.
- The frequency of self-reactive BCR rearrangements is surprisingly high.Indeed, up to 70% of the early immature B cells in the bone marrow are self-reactive3 and most of them are eliminated by negative selection.
- A fifth mechanism belongs on this list, and it is stronger than the four above because it is the only one that sharpens discrimination beyond what the affinity ratio provides. Immune receptors implement kinetic proofreading: signalling requires completion of \(N\) reversible steps, so signalling probability scales roughly as \[(k/(k+k_{off}))^{N}\] and discrimination sharpens as a power of dwell time rather than linearly in affinity8. B cells are not exempt: affinity discrimination for membrane antigen requires proofreading to dominate serial engagement, with threshold dwell times of order seconds9. The affinity distribution and the functional reactivity distribution are therefore related by a nonlinear transform with its own parameters, and clonal selection acts on the second — which is why a broad affinity distribution is compatible with stringent negative selection. And proofreading is not free: because each step is stochastic, discrimination signal‑to‑noise scales as \[\mathcal{O}(N^{1/2}g^{-1/2})\] with \[g=k_{off}/k\], and in the biologically relevant range reliability decreases with added steps10. There is thus a finite optimum to receptor sharpness — a physical reason for the repertoire to maintain graded, overlapping, polyreactive coverage rather than driving every clone toward a delta function.
Thus, falsifying this concept would require careful modeling and definition of the model’s parameters.