Skip to main content

Command Palette

Search for a command to run...

JSONLLM: Vibed AI Model for JSON

I experimented with something called JSONLLM

Updated
•9 min read•View as Markdown
D
CEO of Delino. I created SWC, an open-source compiler project, and previously worked at Deno and Vercel.

I want a fast, capable API for generating JSON. When I build AI-powered apps, I constantly need to generate typed data. And I keep running into the same problems: the output does not quite match the format I need, or it takes too long to arrive.

That led me to build a small proof of concept called JSONLLM. I am not planning to build a business around serving this model. What I really want is for a company like OpenAI or Anthropic to solve this problem properly. But I cannot expect them to build something just because I ask. So I decided to run an experiment myself and prepare the results for release.

I am a developer, and I do not specialize in AI research. I have tried all sorts of things because I am curious about how far LLMs can go. With current coding models, I could even get a training experiment running through vibe coding. I also made plenty of mistakes along the way.

The experiment showed a benefit from batching independent fields. It also exposed substantial problems with accuracy, long responses, and concurrent requests. I want to share both the useful results and the limits.

The problem I kept encountering was structured output. More specifically, I mean giving a model a JSON schema and asking it to produce data that conforms to that schema. Even with Structured Outputs available, the generation felt too slow for the user experiences I wanted to build. As the output grew, the delay became a serious problem.

JSON is verbose. When a model generates an entire JSON object as text, it emits field names and structural tokens alongside the actual values. Structure that the application already knows becomes part of the generation process. I kept wondering how much of that work we could avoid.

The main application I explored in this POC was generative UI, or GenUI. I can see clear uses for it in shopping apps and personal assistants, and I think it could become an important way of building user experiences. But when an interface is represented as JSON generated by a model, JSON generation speed becomes a direct constraint on the experience. If the data needed to build the interface arrives late, the user waits.

While thinking about this, I came across an experiment combining json-render with Jev. I had already been considering the JSON generation problem, but that post gave me a more concrete idea of how to approach it. I wanted something that could go beyond component selection and quickly generate property values and JSON trees.

I also drew heavily on TypeLLM. It was an important reference and a source of ideas. TypeLLM produces typed outputs without changing the underlying model’s architecture or weights. It already supports strings and numbers, parallel execution of independent fields, dependencies between fields, and shared-context reuse. I am not claiming those features as new ideas from JSONLLM.

With JSONLLM, I experimented with training a model for this execution interface. I fine-tuned Qwen3.5-4B with LoRA on field-level decisions and value generation, then connected it to a custom runtime.

The application supplies the structure and rules. The model answers questions about values that have not yet been determined. Code handles calculations, exact copies of existing data, state transitions, and final JSON assembly. Independent fields run together; fields that need earlier answers run in dependency order. I also implemented shared-context processing, where the runtime processes common context once and copies its cache for the field branches. Design documentation

The repository’s order example makes this more concrete. A user says they want to change their shipping address and that it is not urgent. The model decides which component to use and whether the request is urgent. Code retrieves the order details, multiplies the unit price by the quantity, and assembles the final UI data.

The current implementation requires the application to define the available components and structure. It does not freely design arbitrary JSON trees or generate general workflows. There is still a significant gap between this POC and what I ultimately want.

I used Qwen3.5-4B because this was a personal experiment with a limited budget. I was not planning to make money from it, so spending heavily on a larger model from the start was difficult to justify.

The dataset contained 8,000 training records, 1,000 validation records, and 1,000 test records, split evenly between English and Korean. A DeepSeek model generated the synthetic data. After converting the training records into 15,619 field-level examples, one epoch of training took about 102 minutes. Model card

The overall process cost more than it needed to. Looking back, I could have done it much more cheaply. My lack of experience with AI training led to some inefficient decisions. I created multiple working repositories and trained multiple times. Early on, I also failed to check the applicable terms properly and had to repeat training. The model prepared for release was retrained after checking those terms.

I spent about $24 in the first working repository. The second repository’s costs broke down as follows:

Work Cost
Initial training and speed reproduction experiments $140
Shared-context model training and evaluation $232
Latency optimization A/B tests $22.50
Data synthesis, retraining, and evaluation $29.07
Total $423.57

The later comparative benchmark added approximately $64.44. Including the earlier experiments, that brings the overall spending to roughly $512. The additional benchmark estimate includes setup, failed attempts, storage, and recovery of the results.

The $29.07 recorded for the final data synthesis, retraining, and evaluation stage should therefore not be mistaken for the cost of the entire project.

Some of that spending felt avoidable. But I wanted to follow the experiment through and find out what actually worked. The repository I am preparing for release is the cleaned-up version of that process.

For the comparative benchmark, I froze the model weights and tested different execution methods on fresh synthetic tasks. I compared whole-JSON generation, serial field generation, batched independent fields, shared-context fields, and JSON generation through vLLM. Both the original Qwen model and the fine-tuned JSONLLM model ran on the same H100 80 GB. All 540 planned trials completed, with no retraining during this study. Methods and full results

The clearest positive result was batching independent fields. Comparing the same records, batching produced a median latency speedup of about 3.14× over processing fields one at a time, under both models. Accuracy stayed the same for the original model, while JSONLLM answered one additional record correctly out of 256. On this sample, batching improved speed without an observed loss of accuracy.

Shared context produced more mixed results. On the short-context core tasks, it was slower than ordinary batching. It reduced repeated context processing, but also required copying cache state.

In the exploratory longer-context tests, sharing became faster. For JSONLLM, it was about 1.46× faster than ordinary batching at roughly 513 common-context tokens, and about 1.64× faster at 1,026 tokens. However, each condition contained only 32 distinct records. Those results need confirmation on a larger sample.

Comparing field execution with whole-JSON generation exposed an accuracy problem. These are the core results for JSONLLM. Latency and throughput values are medians of the statistics from five trials. A record counts as correct only when every field is correct.

Method Exact accuracy p50 latency p95 latency Correct complete records/sec
Whole JSON 78.52% 2,262 ms 3,816 ms 0.308
Serial fields 64.84% 595 ms 7,376 ms 0.297
Batched fields 65.23% 191 ms 4,913 ms 0.499
Shared-context fields 65.23% 250 ms 4,900 ms 0.490
vLLM JSON 75.39% 323 ms 851 ms 1.781

Shared-context execution had much lower median latency than whole-JSON generation. But its accuracy was approximately 13.3 percentage points lower, and its p95 latency was worse. These results do not support a claim that it generated equally good JSON faster.

vLLM achieved the highest correct throughput and the lowest p95 latency in this comparison. It also handled increased concurrency better than the current JSONLLM runtime. JSONLLM still has a record-level lock: batching fields within a request does not translate into efficient execution of multiple requests.

The tree and workflow results were worse. Even in bounded tasks that asked the model to select the existence, type, and connections of eight candidate nodes, the typed-field methods scored 0% exact-record accuracy. Producing values with valid types did not ensure a correct structure.

The high accuracy from the earlier, restricted evaluation did not carry over to these new tasks. This study provides evidence for batching independent fields, but it does not establish that the current model and runtime solve general JSON generation or GenUI. I have not run a direct performance comparison against TypeLLM either.

I still think the problem matters because I need this capability in the apps I am building now. It applies to GenUI, workflows, and ordinary tree-shaped data. Being able to obtain typed values quickly and accurately is useful across all of them.

I do not know whether my design is the right one. The experiment produced both useful results and failures. I do think it is worth exploring how the model and runtime can work together: which decisions require the model, which operations belong in code, and which independent decisions can run together.

I do not plan to operate this as a service myself. I have limited capital for model serving, and building and serving a model with Luna-level intelligence is not something I can easily do.

I would like OpenAI or Anthropic to build a capable model and API for fast structured-data generation—whether through distillation from a stronger model, improvements to the execution interface, or another approach. Ideally, I could use it through OpenRouter, Vercel AI Gateway, or Cloudflare AI Gateway.

I have organized the code on GitHub, and the model weights, training data, and evaluation materials on Hugging Face. The code, derivative model, and included synthetic data are released under Apache-2.0.

I hope this experiment gets enough attention that people with the resources to build something better take an interest. That may be an unusual reason to train a model, but I really do need this API.