I work on world models, representation learning, graph learning and AI safety.

My research asks what structural assumptions learning systems quietly hard-code, and what becomes possible when those assumptions are made explicit, tested, or removed. I work end to end: the question, the mechanism, experiments, falsification, the manuscript.

Alongside research I founded Varaigen, where I design and operate production AI systems.

publications & manuscriptstwo

public preprint · arxiv:2605.256622026 Can useful graph learning be performed without gradient training? Closed-form graph predictors that match the best measured vanilla GNN on 9 of 9 datasets, and let a deleted label, edge, node or subgraph be re-solved exactly rather than approximately forgotten. Exactness verified across 109 configurations; localized updates 21–45× faster than a full re-solve. manuscriptai safety Which apparent emergent-misalignment effects survive a change of evaluator? A preregistered audit of the measurement itself: two model families, seed-matched organism and control pairs, held-out prompts, automated judges and blinded human raters — agreement on 86 of 90 rows. Some effects survive; others are properties of a narrow task.

graph learningone

ongoing · public evidence repositorygraph learning Is the observed graph the right circuit for inference, or only the evidence about it? Graph models use the observed graph twice — as evidence about the problem and as the circuit inference runs along. Separating the two gives an 11 pp circuit effect against 0.9 pp for repairing the messages.

world modelsone

ongoing, privateworld models When is representation collapse actually harmful? A world-model representation should preserve distinctions by their future consequences, not by generic numerical diversity. Under visual distribution shift it recovers task-relevant latent state at R² 0.96, against 0.21 for a matched baseline.

research infrastructureone

private infrastructureprocess What does research infrastructure look like when it refuses retrospective storytelling? Vyasa keeps hypotheses, preregistrations, evidence ledgers, negative results and research state continuous across sessions, so a programme cannot quietly become a story told afterwards.

record

apr 2025 — nowml lab, iiit hyderabad Undergraduate researcher Lead a small student research team across world models, representation learning, graph learning and AI safety — from hypothesis and theoretical framing through preregistered experiments, falsification and manuscript preparation.
nowvaraigen ↗ Founder Founded Varaigen to design and operate production AI systems: product evaluation, applied research, AI infrastructure and end-to-end system development.
nov 2022 — jul 2024varaisys pvt. ltd. Machine learning researcher IC-ATA mapping and validation workflows, dataset-ingestion systems, configuration-driven training for span and text categorization, model versioning and deployment, evaluation dashboards.
apr 2021 — may 2022mityung Software developer Medzinc, an automated medical-record dashboarding product: ETL ingestion and normalization, role-based access control, backend views, audit logging and REST APIs.

engineering

Varaigen CTEM A commercial-grade continuous threat exposure management platform, built around our patented graph-based approach and sold commercially. Company-owned, so details stay off the page. ctem · graph-based · patented · commercial BUYSELLRENT A campus-gated marketplace: auth, role-based access, listings, search, cart and orders, real-time chat, OTP delivery confirmation. react · node · mongodb · websocket Asset Management System The institutional asset lifecycle for IIIT Hyderabad — procurement to scrapping — with role-based views, audit trails and utilization reports. php · mysql C-Shell A POSIX shell: pipes, redirection, job control, 20+ commands, /proc and TCP built-ins, signal-safe logging and command replay. c · posix · signals Network File System Naming, storage and client processes with path caching, asynchronous writes and networked command handling. c · tcp sockets · lru cache

elsewhere

B.Tech. Honours in Computer Science and Engineering at IIIT Hyderabad, 2023 to 2027, in the Machine Learning Lab. Before research: JEE Main 2022, All India Rank 1790, and nationally ranked junior tennis — AITA U-14 #66. There is a court in the other version of this site.

Working with PyTorch, PyTorch Geometric, DGL, scikit-learn, Hugging Face, Weights & Biases and Hydra. Python, C, C++, Rust, SQL. Linux, Docker, FastAPI, PostgreSQL. Preregistration, ablation design, falsification, reproducibility, human evaluation.

contact

tell me what you are working on.

I am interested in research internships, collaborations and research-engineering roles in world models, representation learning, structured inference and AI safety.

aditya.gaur@students.iiit.ac.in

phone number, home address and private repository links are deliberately not published

Home / Research

Research

Public artifacts are separated from ongoing work on purpose. Nothing here is labelled as accepted until it is.

Interests

Five areas

  • World models and self-supervised representation learning
  • Graph learning and structured inference
  • Exact machine unlearning and auditable learning systems
  • AI safety, model behavior, and evaluation validity
  • Autonomous research infrastructure and empirical methodology
Method

How I approach research

I prefer research questions where the field may be optimizing the wrong object. My work emphasizes explicit hypotheses, preregistered tests, negative evidence, reproducible artifacts, and clearly stated failure boundaries.
Longer account
Home / Publications

Publications

One public artifact, and three manuscripts that are labelled as ongoing because that is what they are.

Public

Peer-reviewable preprint

Closed-Form Node Classification with Exact Graph Unlearning
Aditya Gaur, Charu Sharma
arXiv preprint2026arXiv:2605.25662
Paper Project page Code — on release
Not yet published

Ongoing research and manuscripts

Manuscripts under review are listed as ongoing, never as accepted publications.

Home / Projects

Engineering projects

Systems and product work, kept separate from the research.

Featured
varaigen · commercial ctem platform
Commercial CTEM Platform At Varaigen we built and sold a full commercial-grade Continuous Threat Exposure Management platform, designed around our patented graph-based approach to exposure modeling. The platform is company-owned commercial work, so the specifics — architecture, methods, customers — stay off this page.
CTEMGraph-basedPatented approachShipped & sold
buysellrent · campus marketplace
BUYSELLRENT Built a campus-gated marketplace with authentication, role-based access control, image-based listings, categories, search, filters, cart and order flows, and secure uploads. Added real-time chat, live updates, OTP-based delivery confirmation, rate limiting, and an AI-assisted interaction layer.
ReactNode.jsMongoDBJWTWebSocket
Marketplace · listing · chat · order
asset management system · iiit hyderabad
Asset Management System Designed the database schema and the complete institutional asset lifecycle, covering procurement, assignment, maintenance, utilization and scrapping. Implemented role-based views, audit trails, custom scrapping workflows, and utilization reports for administrators.
PHPMySQLBootstrapSessions
Dashboard · lifecycle · roles · reports
More projects
C-Shell A POSIX-compatible shell supporting pipes, redirection, job control, foreground and background execution, and more than 20 commands. Added custom commands using /proc and TCP, along with signal-safe logging and command replay.
CPOSIXSignalsNetworkingLinux
Network File System Separate naming, storage and client processes supporting read, write, create, delete, list and stream operations. Implemented path caching, asynchronous writes, networked command handling, and command-line tooling.
CTCP socketsLRU cache
Home / Experience

Experience

Download CV
Apr 2025 —
Machine Learning Lab, IIIT Hyderabad
Honours student / Undergraduate researcher
Lead a small student research team across world models, representation learning, graph learning, and AI safety. Drive projects from hypothesis formation and theoretical framing through implementation, preregistered experiments, falsification, reproducibility, and manuscript preparation.
Build research infrastructure using PyTorch, PyTorch Geometric, DGL, Weights & Biases, Hydra, and custom experiment orchestration systems.
Present
Varaigen
Founder
Founded Varaigen to design and operate production AI systems. Work spans product evaluation, applied research, AI infrastructure, and end-to-end system development.
Engineering intelligence, end to end. varaigen.com
Nov 2022 — Jul 2024
VARAISYS Pvt. Ltd.
Machine learning researcher
Built IC-ATA mapping and validation workflows, dataset-ingestion systems, and configuration-driven training pipelines for span and text categorization.
Developed model versioning, online and offline deployment, auto-annotation, suggestion systems, evaluation dashboards, and precision, recall and F1 reporting using Python, spaCy, scikit-learn and REST APIs.
Apr 2021 — May 2022
MITYUNG
Software developer
Worked on Medzinc, an automated medical-record dashboarding product. Built ETL ingestion and normalization, role-based access control, backend views, audit logging, tests and REST APIs using Flask and SQL.
Collaborated with clinicians and product stakeholders to translate operational workflows into software requirements.
Skills
Research World models, representation learning, graph learning, machine unlearning, model-behavior evaluation, experimental design, preregistration, ablation design, reproducibility, human evaluation
Machine learning PyTorch, PyTorch Geometric, DGL, scikit-learn, Hugging Face, pandas, NumPy, Weights & Biases, Hydra
Languages Python, C, C++, Rust, SQL, Bash
Systems and engineering Linux, Git, Docker, FastAPI, Flask, PostgreSQL, MySQL, REST APIs, WebSockets
Education
International Institute of Information Technology, Hyderabad B.Tech. Honours in Computer Science and Engineering Jul 2023 — May 2027 · Machine Learning Lab
Outside research
  • JEE Main 2022 — All India Rank 1790
  • Former nationally ranked junior tennis player — AITA U-14 rank 66
Home / About

About

I am an undergraduate researcher at the Machine Learning Lab at IIIT Hyderabad, where I study world models, representation learning, graph learning, machine unlearning, and AI safety.

My research is centered on a recurring question: what assumptions are our learning systems treating as fixed, even when they should themselves be objects of inference or design?

In graph learning, this led me to investigate how much node-classification performance can be recovered using deterministic closed-form solvers, and what exact guarantees become possible when prediction is an explicit system rather than the result of iterative gradient training. The resulting work develops closed-form predictors with exact retrain-equivalent unlearning for graph-object modifications.

In a separate line of work, I study the distinction between a graph as evidence and a graph as an inference circuit. Most graph models assume that the observed graph should also determine the path along which computation flows. My work asks when that assumption is valid and what becomes possible when the two roles are separated.

My current world-model research studies representation collapse from the perspective of future consequences. Instead of requiring representations to remain generically diverse, I investigate whether they can compress aggressively while preserving precisely the distinctions needed for prediction, action, and control.

I also work on AI-safety evaluation, particularly the measurement of emergent misalignment and the degree to which reported effects survive changes in model family, intervention, behavioral construction, evaluator, and human framing.

Across these projects, I emphasize preregistration, reproducibility, negative results, and experiments designed to falsify the central claim. I have also built private research infrastructure for maintaining long-running autonomous research programs across hypotheses, experiments, evidence, and manuscript development.

Alongside academic research, I founded Varaigen, where I work on production AI systems and applied research.

Aditya Gaur
IIIT Hyderabad · Machine Learning Lab Hyderabad, India
Now

What I am working on now

Collapse-contained world models Testing whether future-consequence equivalence can replace generic anti-collapse objectives in JEPA-style learning.
Inference structure in graph learning Studying when the computational circuit should be separated from the observed evidence graph.
Human-grounded AI evaluation Auditing which model-behavior effects survive evaluator changes and blinded human judgment.
Autonomous research infrastructure Developing tools for persistent, falsification-oriented research campaigns.
For organisers and applications

Third-person bio

copy and paste

Aditya Gaur is an undergraduate researcher at the Machine Learning Lab at IIIT Hyderabad and the founder of Varaigen. His work spans world models, representation learning, graph learning, exact machine unlearning, AI-safety evaluation, and autonomous research infrastructure. He is the first author of “Closed-Form Node Classification with Exact Graph Unlearning,” which develops deterministic graph predictors with exact retrain-equivalent unlearning. His current research studies decision-grounded representation collapse, the separation of graphs as evidence from graphs as inference circuits, and human-grounded evaluation of emergent misalignment.

Research / Exact graph unlearning
Public preprint, 2026 arXiv:2605.25662

Closed-Form Node Classification with Exact Graph Unlearning

Deterministic closed-form graph predictors that recover competitive GNN performance while enabling exact, retrain-equivalent deletion of graph objects.

Read on arXiv PDF Code — on release
9 / 9Datasets matching or exceeding the best measured vanilla GCN, GraphSAGE or GAT
109Configurations with verified exact graph-object unlearning
21–45×Faster localized updates than full re-solving on ogbn-arxiv
14Graph benchmarks evaluated
The question

Can useful graph learning be performed without gradient training — and if the predictor is an explicit solution rather than a trained network, what guarantees come for free?

Why closed-form graph learning

Node classification is normally solved by iterative gradient training over a message-passing network. That choice is convenient, but it makes deletion hard: once a label, feature, edge, node or subgraph has influenced the optimizer's trajectory, removing its effect means either retraining or accepting an approximation.

A predictor that is an explicit solution of a deterministic system does not have that problem. The question is whether such a predictor is still competitive.

Core hypothesis

Competitive node classification can be recovered by routing between two closed-form predictors chosen by the graph's structure, and because both are explicit solutions of deterministic systems, modification becomes a re-solve rather than an approximate forgetting procedure.

Routed architecture

On assortative graphs the framework uses SGC-style propagation with Ridge regression. On heterophilous graphs it uses a layer-wise closed-form feature-refinement network. A route is selected per graph, so the framework does not force one inductive bias onto both regimes.

FigureRouted architecture and localized unlearning
The routed framework: an input graph is scored for adjusted homophily, then routed to Pipeline A (propagation plus Ridge solve) when assortative or Pipeline B (multi-scale features, LCF-Net Ridge layers, Gaussian KRR head) when heterophilous. Both pipelines end in deterministic functions of the training data, so a deletion is re-solved exactly.
The router reads adjusted homophily once, then commits to one of two closed-form pipelines. Because both ends are explicit solutions rather than trained weights, a deleted label, edge, node or subgraph is re-solved exactly — Theorem 1 bounds the work to the deleted object's L-hop neighbourhood.
Exact graph-object unlearning

Because the resulting predictors are explicit solutions of deterministic systems, modified labels, features, edges, nodes and subgraphs can be re-solved exactly rather than approximately forgotten. Exactness was verified across 109 configurations.

Experimental design

Fourteen graph benchmarks, spanning assortative and heterophilous regimes. Accuracy is compared against the best measured vanilla two-layer GCN, GraphSAGE or GAT result. Unlearning is evaluated as exact re-solve equivalence across label, feature, edge, node and subgraph modifications. Locality and update cost are measured on ogbn-arxiv.

Main result

The routed closed-form framework matched or exceeded the best measured vanilla two-layer GCN, GraphSAGE or GAT result on 9 of 9 datasets, while supporting exact deletion that gradient-trained baselines cannot offer.

Locality and efficiency

K-hop locality is proved for the Ridge components, so a modification only requires re-solving within its neighbourhood. Localized updates ran 21 to 45 times faster than full re-solving on ogbn-arxiv, and approximately one million times faster than gradient retraining in the reported comparison.

Privacy boundary

Exactness here means retrain equivalence for the specified graph-object modifications: the updated predictor is the one you would have obtained had the deleted object never been present. That is a deletion guarantee, not a differential-privacy guarantee, and it is stated as such in the paper.

Reproducibility

Because the predictors are deterministic solutions, there is no optimizer schedule or seed to match: given the same graph and the same route, the solution is the solution. The intended public code repository will be linked here on release.

Citation
Closed-Form Node Classification with Exact Graph Unlearning
Aditya Gaur, Charu Sharma
arXiv preprint2026arXiv:2605.25662
Paper
What comes next

The public code repository is the immediate next step, so that the exactness claims can be re-run rather than taken on trust.

Research / Decision-grounded world models
Ongoing private research Manuscript and additional results available privately

Decision-Grounded Representations for Predictive World Models

A world-model representation should preserve distinctions according to their future consequences, not merely maintain generic numerical diversity.

R² 0.96Task-relevant latent state recovered, controlled visual-OOD setting
R² 0.21Matched baseline, same setting
Observation A · Observation B Different pixels, same future response. Safe to merge
Observation C · Observation D Similar pixels, different future response. Must remain separate
The question

Can a representation compress aggressively while preserving exactly the distinctions needed for future prediction and control?

Why the existing formulation may be insufficient

Most anti-collapse methods define collapse geometrically through variance, covariance, rank, or teacher-target constraints. Those objectives measure the shape of the latent space, not what the latent space is for.

A representation can be numerically diverse and still erase the only distinction that matters for future behavior. It can also be highly compressed while preserving everything required for prediction and control.

Core hypothesis

Anti-collapse objectives should preserve response-relevant equivalence classes rather than arbitrary latent diversity. Collapse is only harmful when it merges observations whose futures differ.

Method

The mechanism is not public yet. What can be said publicly is the criterion it optimizes: distinctions are retained or discarded according to their effect on future prediction and control, rather than according to a geometric diversity term.

Manuscript and mechanism available privately on request.

Experimental design

The hypothesis is tested across controlled latent environments, visual distribution shift, raw pixels, dependent observations, and decision-native offline reinforcement-learning tasks. Controlled settings come first because they make the claim falsifiable before it is scaled.

Main result

In a controlled visual-OOD setting, the decision-grounded representation recovered task-relevant latent state at approximately R² 0.96, compared with approximately 0.21 for a matched baseline. The project has since transferred to raw pixels.

Failure boundaries

Current work focuses on raw-pixel transfer, correlated evidence, richer environments, and identifying the boundary at which the mechanism fails. A result that only holds in the setting it was designed for is not yet a result.

Status and disclosure

Ongoing private research. One headline result is published here; the mechanism, the full experimental ledger and the manuscript are shared privately rather than posted.

Research / Evidence, not circuits
Ongoing research with a public evidence repository

Graphs Are Evidence, Not Circuits

The observed graph may describe the evidence available to a learning problem without being the correct computational circuit for inference.

Evidence graph — the question What was recorded: interactions, measurements, similarity, communication, incomplete observation.
Inference circuit — a separate object The path computation should actually take. It may need to be designed or learned, not inherited.
11 ppCircuit effect — changing the inference circuit
0.9 ppMessage-repair effect — improving messages over the observed graph
The question

Is the observed graph the right circuit for inference, or only the evidence about it?

Why the existing formulation may be insufficient

A message-passing GNN does not merely learn parameters over a graph. It also hard-codes a hypothesis about the structure of inference: information should flow along the observed edges.

That assumption is natural when the observed graph is itself the factorization of the inference problem. It is not generally justified when the graph only records interactions, measurements, similarity, communication, or incomplete evidence.

Core hypothesis

The project treats the observed graph as the question and the computational circuit as a separate object. If the two roles are genuinely distinct, then changing the circuit should matter more than repairing the messages sent along the observed one.

Method and instruments

Architectures and diagnostics that let the circuit be varied independently of the evidence, so the two effects can be measured separately instead of being confounded inside one model.

Experimental design

Every experiment registers a prospective prediction before it runs, and the outcome is recorded whether or not it matches. Preregistrations, result artifacts and audits are frozen in the public repository rather than summarized after the fact.

Main result

Changing the inference circuit produced an 11 percentage-point effect; repairing the messages propagated over the observed graph produced 0.9. Compiling the circuit rather than inheriting it improved on that again — the exact artifact lives in the repository's frozen results.

Negative results and audits

Falsified hypotheses are kept, not deleted. The repository retains negative results alongside theorem and circuit audits, so a reader can see which versions of the claim did not survive.

Reproducibility
  • Frozen preregistrations and result artifacts
  • Negative results retained, not discarded
  • Theorem and circuit audits
  • Reusable packages and reproducibility scripts
What comes next

The paper is not public yet. The evidence repository is, and it is the honest version of the claim: implementations, preregistrations, negative results and audits, linked here as soon as the repository is ready to be read by strangers.

Research / Emergent-misalignment audit
Manuscript, 2026 Additional materials available privately

Human-Grounded Audits of Emergent Misalignment

A preregistered audit of which apparent emergent-misalignment effects survive changes in model family, intervention, behavioral construction, evaluator, and human framing.

01 Intervention Seed-matched organism and control pairs
02 Construction How the behavior is elicited and held out
03 Judge Automated verdicts on held-out prompts
04 Humans Blinded human raters, same rows
86 / 90Rows where blinded human and automated binary verdicts agree
2Model families, five seed-matched organism and control pairs
The question

Which apparent emergent-misalignment effects survive a change of evaluator?

FigureWhat survives the audit, in six steps
Six numbered steps. Swapping the judge on the same rows drops the catch rate to 3-44 percent. Five seed-matched pairs one KL knob apart replicate the effect: organisms plus 13 points, controls exactly on the zero line. The mean difference installs on a fresh model at plus .65. Recompressing after subtracting it shows rank 1 loses the skill (.77) while rank 2 keeps it (.03). Blinded humans re-reading the clean result drift to .14 by hand against a .10 bar.
The same rows, measured six ways. Steps 1–3 ask whether the reported effect is a property of the model or of the evaluation package; steps 4–6 install the difference on a fresh model and hand the cleaned result back to blinded human raters, who read it slightly higher than the automated judge does.
Why the existing formulation may be insufficient

Claims about emergent misalignment are only as strong as the evaluation package used to measure them. If a reported effect depends on one intervention, one behavioral construction and one automated judge, then the finding may be a property of the measurement rather than of the model.

Core hypothesis

Apparent transport of emergent misalignment is sensitive to the evaluation package. Some effects will survive a change of model family, construction and evaluator; others will not, and the difference is measurable.

Instrument

The audit stress-tests the measurement process itself rather than one model: two model families, five seed-matched organism and control pairs, preregistered checkpoints, held-out prompts, automated judges, and blinded human raters scoring the same rows.

Main result

Apparent transport depends strongly on the intervention, the behavioral construction, and who is doing the measurement. Some effects survive broad human evaluation, while others are properties of a narrow task or evaluator. Blinded human and automated binary verdicts agreed on 86 of 90 rows, so the disagreement is not simply judge noise.

Boundaries

The audit separates broad behavioral transport from narrow task-specific estimates. Where an effect is narrow, the paper says so rather than reporting the larger number.

Status and disclosure

Manuscript, 2026. Preregistrations, the full evaluation package and human-rating materials are shared privately while the work is in review.

Research / Vyasa
Private research infrastructure

Vyasa: Autonomous Research Harness

A research harness for turning open-ended ideas into persistent, falsification-oriented experimental programs.

Question Prospective prediction Execution Falsification Synthesis Next step
The question

What does research infrastructure look like when it refuses retrospective storytelling?

Why it exists

Vyasa is private infrastructure I developed to support long-running research programs across world models, graph learning, and AI safety.

It maintains structured hypotheses, experiment registries, preregistrations, evidence ledgers, negative results, decision rules, and research-state continuity across agent sessions. The system is designed to prevent research from degenerating into disconnected experiments or retrospective storytelling.

Capabilities
  • Persistent research state
  • Hypothesis and finding registries
  • Preregistered experiment gates
  • Negative-result retention
  • Evidence and claim ledgers
  • Reproducibility tracking
  • Autonomous experiment planning
  • Multi-project research vaults
  • Human-readable research summaries
The loop it enforces

It supports repeated cycles of question formation, prospective prediction, experiment execution, falsification, synthesis, and next-step selection — in that order, with the prediction written down before the run.

Status

Private research infrastructure. The repository is not public; the design and its rationale can be discussed.

BibTeX copied
Bio copied