Six FLock.io academic papers on decentralised AI have been accepted to the prestigious CIKM 2026 Full Research Track, a globally top-tier academic conference in information retrieval, knowledge management and database systems.
Each of the six papers tackles a core challenge in AI and strengthens the foundations of decentralised AI, from training efficiency to data privacy. The underlying theme of all of these six papers is graph-structured probabilistic and topological reasoning for LLM systems. All were authored by Dr Zehua Cheng, FLock.io’s Chief AI Scientist.
The 35th Conference on Information and Knowledge Management (CIKM) will be held 7–11 November 2026 in Rome, Italy. Competition is fierce at CIKM, with acceptance rates typically ranging between 15% and 30%.
[ 👋 Hi there! If you’re here to find out more about FLock, follow us on X and email us at hello@flock.io to learn what we do in helping companies meet compliance in AI.]
Rather than treating language models, model weights or training data as flat sequences or monolithic vectors, every paper explicitly converts complex internal structures into probabilistic factor graphs or topological dependency networks. They apply classical graphical algorithms (such as loopy belief propagation, Bethe free energy minimisation, cavity methods, graph coloring and topological graph knock-outs).
They focus on these key areas:
Data & privacy: They decouple identity from logical structure using entity-coreference graphs.
Inference & decoding: They repair in-context learning, resolve structured output dependencies, and quantify spatial uncertainty via attention- and mutual-information-derived factor graphs.
Training & optimization: They optimise Transformer pre-training schedules and predict weight-space model mergeability through online curvature and interaction graphs.
The CIKM is a premier international forum for the presentation and discussion of research in information retrieval, knowledge management, data mining and database systems. It brings together researchers and practitioners from academia and industry for its mission to address the most challenging problems in the design and development of future information and knowledge systems.
The conference fosters high-quality contributions spanning theoretical models, algorithmic advances, system architectures, experimental evaluation and real-world applications.
The six papers:
- “Graph-Driven Context-Preserving Anonymization for LLM Scaling”
- “Bethe Consensus Decoding: Joint Inference on Empirical Pairwise Factor Graphs for Structured LLM Outputs”
- “Curvature Graph Mergeability: Predicting Model Compatibility from Per-Layer Hessian Geometry”
- “In-Context Learning as Approximate Belief Propagation”
- “Gradient Interaction Graph Optimization”
- “Cavity-Calibrated Prediction Sets: Single-Pass Conformal Inference for Structured Outputs Over Graphical Models”
Graph-Driven Context-Preserving Anonymisation
Some of the most valuable AI training data contains sensitive personal information, making it difficult to use without exposing people.
This approach removes identifying information while preserving the meaning and reasoning in the original data. On the benchmark, downstream accuracy stayed within 0.1 percentage points of the original, showing that private data can remain useful after anonymisation.
Bethe Consensus Decoding (BCD)
Ask an LLM the same question several times, and you can vote on the most common answer. The problem is that if the answer has multiple fields, choosing each field independently can create contradictions, such as pairing “junior” experience with a $220k salary.
BCD looks at all the answers together and finds a combination that makes sense as a whole, improving exact-match accuracy by 4.6 points over majority voting without additional training.
Curvature Graph Mergeability (CGM)
What if you want to combine two separately trained AI models? First, you need to know whether they will actually work well together, but finding that out can be slow and expensive.
CGM predicts whether two models can be merged using a few cheap measurements, making the process roughly 14,500x cheaper than the standard approach while also identifying which parts of the models could cause problems.
In-Context Learning as Belief Propagation (ICL-BP)
Give an LLM a few examples to follow, and one misleading example can derail the answer.
ICL-BP helps the model identify which examples are actually useful and reduces the influence of unreliable ones. Across eight benchmarks, this improves median accuracy by 7.4 points without requiring the model to be retrained.
Gradient Interaction Graph Optimization (GIGO)
Training a model means adjusting its internal settings again and again to reduce mistakes. The problem is that different parts of the model often work with slightly outdated information from one another, which wastes training steps.
GIGO maps how these parts influence each other and updates them in a smarter order, reaching the same GPT-2 quality with 32–34% fewer training steps while working with existing optimisers.
Cavity-Calibrated Prediction Sets
In high-stakes AI, getting an answer isn’t enough because you also need to know how confident the model should be. For complex predictions, standard methods can either make the model look more certain than it is or produce such broad answers that they become difficult to use.
CCPS provides mathematically backed, calibrated uncertainty estimates in a single pass, giving users a more reliable sense of where the model's prediction stands.
More about FLock.io
FLock.io is an AI research and infrastructure company pioneering enterprise-grade federated learning and distributed AI solutions. Its decentralised federated learning architecture and production-ready platforms (AI Arena, FL Alliance, and FLock API Platform) enable organisations to train and deploy their own custom AI models on local hardware while maintaining full data privacy, model ownership, and regulatory alignment by design.
FLock.io is internationally recognised for its academic research, including the NeurIPS award-winning paper “FLock: Defending Malicious Behaviors in Federated Learning with Blockchain” and sponsors computer science PhD students at the University of Oxford.






