Detecting a false common space, why L2 is not an ontology.

Mapping mixed-norm aggregations - notes from a long conversation about a shared vocabulary with biotech.


An adventure in itself. What we temporarily arrived at is the simplest description of what the tool actually returns.

In any calculation it is natural to discard data whose influence is negligible. We call this thresholds, filters, cutoffs - depending on where they sit and what they are for. A banal example is the rigid, hand-applied reflex of computing interactions with a ball:

The ball: from far away it is a point (a monopole / a center of mass). As we approach, moments appear, then a surface, then an interior. Nonlocal contributions are damped by a cascade (multipoles, RG, coarse-graining) until they fall below the resolution of the calculation and become "zero or noise". Relations are computed locally; the rest is statistics or is ignored. Simple. Everyone does it.

HMFoDG simply computes those damping levels automatically, in a loop, so that the context (the granularity of the data) decides how much of a given contribution to keep. And that, in the context of the result, is everything. For a ball (or a roughly radial cluster) it is obvious. For more complicated structures there was no clean way to bite it. One introduces a grid, a lattice. But then, in a nearly empty space (low density of arguments, large distances - the standard molecular regime, where most of the volume is void) addressing becomes awkward, and every argument sits at a different damping level, which means it is counted under a completely different influence rule. Isolated tools do something about this. None of them integrate it, in a loop, into a local influence level and immediately return a metric for that calculation. HMFoDG does this almost by accident: that was not the goal. I was only trying to tidy something up, and it turned out that all the pieces of the puzzle already existed.

Getting to that conclusion was the adventure. Here it is. In square brackets I note what we established against a checklist: which mathematical apparatuses the people in the room already know (have in their curriculum) and in what form - because that is where the communication friction sat. One does not ordinarily terrorize engineers with a full derivation and a proof of the entire theory behind an apparatus. They have other things to learn, they will not need this, and a lifetime is finite.

We sat down with colleagues from biotech over the HMFoDG loop and asked:

Do you suspect that particular factors in the correlations you use do not have a p = 2 influence on the rest?

(Vectors, L2, p = 2, correlation - they know them, they use them, nothing to explain.)

The suspicion is there, of course. Which norm to apply, and why, is another matter. They cannot justify L2 itself except by “that is how we have always done it.”

Why other p?

(the only piece of theory that actually matters here);

In formulas (already in school-level physics) one keeps meeting cascades of power means. They are in Newton, in special relativity, in Starobinsky, in field theory - everywhere. This is a property of sets that are not abstract curiosities invented for theory (cofinite sets, complements, ultrafilters, pathological Hamel bases). Physics keeps operating the same power-cascade structure because that is how observation builds an observational data set. Observation is homogeneous: one readout, one resolution (hence the cascades are truncated), far fields, static stability, local effectiveness. Observation is a projection onto an x0 (instrument, scale, direction). Data are fibers (the full path, heading, history).

A working physical formula reports a false distance (a projection) - that is, an influence - not the source spacing, and silently damps what the projection cannot distinguish. The cascade of exponents is the logarithm of resolution. Truncating the cascade is honest, because the instrument’s mantissa is truncated too.

Formulas are structurally identical because every finite observer must rescale channels to a common unit (a power), fold them into a scalar (aggregation under a norm), and throw away floors below resolution (cutoff). It is not that nature “likes” Lp.

HMFoDG computes where to put the cutoff and which p yields the least false quasi-norms for the components.

One norm ≠ ontology. We jumped the resistance with a question: do you have a proof that all data live under a single norm? Data are not the world. The way they are collected is what it is. We study data from the world, not the world.

Then we went back to the set and to building a graph. For most engineering professions a graph does not go beyond "a picture, a schematic". Graph + neighborhood - networks people understand widely. Geometry on a graph is another story: they treat it as a set without edges, a sketch for imagination, the school-entry version of graphs. Fair enough; in that field going deeper may be unnecessary. They have other things to learn.

I spent a moment wondering why I see it differently - as a structure that depends on the choice of x0 - and concluded that for me it was also once “a picture.” Years of doing arithmetic on peculiar graphs simply changed how I think.

State (v, γ), paths - chemical intuition works here. The formal notation is readable. Descriptive implications of fuzzy edges as lists of conditions along a trajectory also.

Local entropy + a lambda scan - they understand operationally. Fisher as a formula is acceptable; do not scare them with Amari. Another case where application only needs that someone, somewhere, proved it, and that it works for them. I do the same: I use ready blocks from mathematics. I do not invent new objects or new theories. I analyze what is there. Inventing new objects is too high an intellectual floor for me; I do not reach there. Creative use of other people’s ideas is all I can do.

Prefix / tree / NN = LCP - as a drawing of a little tree (a graph), yes. Do not scare them with the word “ultrametric.” “Little tree” sounds better.

Flow + reprojection - they understand operationally, from a diagram; they have physical analogies.

Lorentz, interpolation, Orlicz - they do not know them. They know that something frightening lives there; not their department. I need those only to check whether a lemma derived in iterations falls apart.

p-adics as theory - as a labeling system, yes. They ask not to be told more deeply why. It is enough that it works.

Noncommutative algebras - it is enough to say that operations may be performed in a given order and in no other. They know perfectly well that a hen requires an egg, and eggs were laid by dinosaurs, so hens are possible, and one must not tinker with the order of operations.

Hessian on the complement, cofinite as a no-go - conditions on which databases may be processed are unnecessary to them. They do not have abstract bases anyway, only ones from real measurements.

So we have several cutoff points that have to be given as kitchen recipes, without frightening anyone with a hundred pages of mid-20th-century analysis that assumes the reader already has the knowledge base.

After a long discussion we reached a conclusion. The point is only this: do not impose one norm from above (customarily the quadratic one). Build a co-occurrence graph (and that they already have in the data: what happened to the object of study, or what was done to it). Fold everything with a power formula Rw. Find the local sharpness of the image (entropy / Fisher at a given scale). Write the hierarchy in a shared code (not a lattice of relations). The path on the graph is part of the state, so the same two nodes have a different measure of “how many passages there are” and a different length. There is no global space. There is a local map, for as long as gluing holds. When it does not hold, that shows up in the returned shared code, which has no position in common with another code (some have none in common with anyone).

This is an application of known apparatuses (power means, a graph with state, a prefix tree, Fisher information as contrast), not a new theory. The notes are only a check of whether one is allowed to compose them this way and whether the composition falls apart (but I am not to frighten anyone about whether the spectrum or the kernel falls apart, or whether a contraction is reached; the notes are a list of admissibility). Fragmentarily, these apparatuses are already used across industries. But outside math/physics nobody keeps them in one circuit:

drop the global space → graph → power aggregation → scan scale with Fisher information → ultrametric address → check that the spectrum / another apparatus does not wreck it → go on.

The apparatuses used are known in biotech; they are simply not named, in the notes, in the language used there. And they are not used as a single consistency filter. Ultrametricity is in UPGMA. Lp is in regularization. The graph is in PPI. Fisher is in MLE. In the notes there is an analytical workshop.

The tool solves the problem of detecting a false common space. And of course it indicates transitions for trajectories where mixed norms have equivalence within the accepted resolution.


The kitchen recipe (what the loop actually does). Do not impose one norm from above (usually the square).

  1. Build a co-occurrence graph from what the object of study underwent or was subjected to. Biotech already has this graph in the data.

  2. Fold local values with a power aggregator
    R_w(y) = (∑ᵢ αᵢ |yᵢ|^w)^{1/w},
    with w = 1 neutral, w < 1 diffuse / generalizing, w > 1 selective / amplifying the dominant.

  3. Measure local sharpness (entropy or Fisher at that scale). That chooses w and the cutoff.

  4. Write the hierarchy as a shared address code (prefixes), not as a global lattice of relations.

  5. Treat the path as part of the state. The same two nodes can have different “how many passages” and different length, depending on the route and the current norm.

  6. If gluing between local maps fails, the returned codes simply do not share a position. That is the detection of a false common space.

This is nothing new. It is one circuit made of apparatuses already used, separately, in the building:

Piece in the notes Already in biotech / ML under another name
Ultrametric address / LCP tree UPGMA, hierarchical clustering
Variable p/w regularization, robust losses, different moments
Graph with state PPI, pathway, event graphs
Fisher as contrast MLE, information about a parameter
Cutoff / horizon LOD, resolution, sparsity threshold

Fragmentarily, every floor is already in use. Outside math/physics they are rarely kept as one consistency filter.

The tool’s job is narrower than it sounds: detect when a shared space is a projection artifact, and mark the transitions where mixed norms are still equivalent within the accepted resolution.


Input data: order that observation can actually support;

The claim is not about all sets. It is about sets that can be data: finite, or locally finite, relations coming from observation. Not cofinite pathologies, not choice on a continuum.

On such a carrier even mixed norms do not destroy order. Every point can be an x0. The neighborhood is finite. R_w returns a scalar. Scalars on the star give a preorder - and, more usefully, an order relative to a projection a human can read.

Most such orders are already good enough in L2. Observation typically folds many weak, weakly dependent contributions. Central limit → Gauss → variance / Fisher / least squares. Kinetic energy starts at v^2. RMS, chi-squared, PCA, cosine on normalized vectors live in that regime.

Other p appear when typicality breaks: a dominant (p → infty), rare events, a heavy lower tail, intermediate factors living on another scale. Physics “works” on the square until the calculation approaches the ball, or the tail gets fat. Then one adds the next term of the same cascade.

Biotech leaves that regime early. Already in chemistry one does not write the physics of everything above small atoms. A lead nucleus is a flops problem, not a logic problem. A “particle” in biotech is a huge number of atoms plus dynamics plus conditions. There is no isolated electron in an empty universe. Even if every micro-piece were quadratic, distances and environments do not reassemble into one global p. Add event dynamics and correlations start looking like ritual: only at the second full moon after a certain conjunction.

Statistics supplies generic norms. The research question is usually the non-generic composite: trajectories of events. A drug does not act the same on everyone. Not everyone gets the disease.

If a projection can be read, it can be locally ordered by a homogeneous scalar. Typically that scalar is quadratic; the other floors are corrections when typicality or resolution lies.

If there is no repeatable trunk, no homogeneous aggregation will stabilize a preorder around x0. Fisher has no peak in the middle of the interval, only a smear at the edge; effective k jumps; LCP between pairs sits near 0 - Pearson near zero with an error bar that could flip the sign.

That can be noise from a larger environment, or unprovable fundamental chance. Locally it looks like snow. Change x0, idiolect, horizon, or prefix depth, and a trunk may appear. A cutoff that is brutal on this scale can vanish on the next. On pure noise with no trunk, search should mark everything with zero LCP.

Careful: no trunk on this tower is not no trunk at all. The grid may be too coarse, or patience too short. A hint of scale. Not a verdict that the world is random.

HMFoDG does not invent a new geometry. It computes, locally, which power and which cutoff make mixed contributions comparable and flags when a shared space is only a projection artifact.

If You are intrested to dig deeper or just need someone to solve analitical assumptions in Your project....

The tool is implemented as a prototype tested for several months. If you are interested in the computational side then the entire loop of the following two equations (aggregation + generalized Fisher with free metric):

  1. Projection onto x0 → radial graph + v_i = x_i - x0.

  2. Local N(v), ordering y_i.

  3. Aggregator: R_w(y) = (∑ α_i |y_i|^w )^{1/w} [w adaptive, w=1 neutral].
    w < 1: diffuse/generalize
    w > 1: selective/amplify dominant

  4. Hierarchical addressing: directional/p-adic prefixes → ultrametric d = p^{-LCP} or e^{-λ LCP}.
    Lemma: NN ≡ max LCP → prefix tree = efficient hierarchy.

  5. Lambda/w from Fisher-like: local entropy H(N(v)), density, or spectral radius embeddability (ρ² < α).

  6. Flow: repeat aggregation → update metric/prefixes → reprojection with new data (redshift horizon).

Yields:

Advantages for ML/embeddings:

  • Deterministic hierarchy (sharp taxonomies, noise control).

  • Ordering of training instead of post-hoc censorship.

  • Strong local metric deformation (tunnels, contrasts).

  • Easily implementable (prefixes + LCP << cosine brute force).

This procedure is applied in the tool, from which I can list embeddings.

The computational and proof explanation of this loop is in:

And curiosities are listed here:

Feel free to contact me if You have an analitical assumptions to solve.

-- Jack Kowalski

Back to top