Engineering the Next Frontier of Artificial Intelligence

Cymela is an independent AI research effort working on continuous latent reasoning, MultiThink architecture, and a terminal coding agent.

Monarch·Neuralese·MultiThink·CLI

EST. 2026 FOUNDED
6.93B CURRENT MODEL SCALE
SPARSE MoE 64 EXPERTS, 8 ACTIVE
INDEPENDENT SELF-FUNDED RESEARCH
Our Thesis

One big bet.
Latent reasoning.

Read about Neuralese →

Cymela's core thesis is continuous latent reasoning: instead of a model thinking only by generating text token-by-token, we're researching how it can carry reasoning forward as a continuous internal state, closer to how the underlying computation actually happens, before it ever gets forced into words.

Our current research checkpoint is Monarch Chrysalis 1, a 6.93B sparse mixture-of-experts model, 64 experts per layer with 8 active per token, extended with our own latent-reasoning modules. It is not derived from Qwen. The earlier checkpoint, Hyper v1, is a 3.10B dense model built on Qwen2.5-3B-Instruct and released under that research-only license; it is the one with public weights. It's early, and the behavior is still inconsistent: the model thinks in latent space and that thinking helps, but it isn't yet specific to the question you asked. That gap is the work.

Alongside the model research, we build and ship Cymela CLI, a terminal coding agent you can install and use right now. One is a research bet; the other already works.

What We're Working On

Four fronts. One thesis.

FIG.01

Latent Reasoning

Our core research thesis: models that carry reasoning forward as continuous internal state instead of only through generated text. Two checkpoints trained, and every result published on measured weights, including the ones that argue against us.

Read about Neuralese →
FIG.02

Cymela CLI

A terminal coding agent, shipped: file editing, search, git, and shell tools, working across multiple model providers today.

Install the CLI →
FIG.03

Agent Orchestration

An early research direction we call MultiThink: two models working at the same time rather than in turn, one generating while the other checks and corrects. Right now this is exploratory, not a shipped system.

See where it stands →
FIG.04

Training Approach

We train across whatever compute we can get: cloud TPU access and personal GPU hardware. Resourceful, not enterprise-scale, yet.

Explore the roadmap →

Think beyond words.

Read about Neuralese
What is Samaritan?