Open research, local inference

Small open models you can measure

Eschatia Labs is an international open-source research lab. We build small language models for decisions and local AI agents, and study explainable deepfake detection.

01 / Model families

Choose a task

Text classification & decisions

JevAlt

Route a support ticket, apply a policy, or choose an action. Three 4B language models return decisions with probabilities through the Jev API.

Karar-4BTurkish · 4B parameters

96.8% held-out accuracy · Kev-4B 87.1%

Deem-4BEnglish · 4B parameters

94.7% held-out accuracy · Kev-4B 84.7%

Wähler-4BGerman · 4B parameters

92.0% held-out accuracy · Kev-4B 81.1%

Held-out tests share the training data pipeline. These scores measure the target decision task. Q4_K_M uses about 3 GB RAM at 4k context with the JevAlt server. Refit calibration on your own data and keep authorization outside the model.

Local AI agents & tool use

Tholos-2B

A 2B model that chooses tools for tables, notes and tasks on your own machine.

137/160 tasks

Tholos-Bench, Q4_K_M, llama.cpp JSON schema, Kaggle T4. MiniCPM5-2B: 112/160. Only Tholos-2B was fine-tuned for this task format.

Q4_K_M file: 1.56 GB. Use Ollama or llama.cpp. Validate tool calls and retain approvals for consequential actions.

Explainable deepfake detection

xdfdet

Eight EfficientNet-B4 detectors with Grad-CAM maps showing which facial regions influence a prediction.

0.8981AUC on FaceForensics++

Best released checkpoint, aug-cutout-black. Baseline: 0.8684. Each checkpoint used its own random split. On 398 unseen DFDC videos, AUC drops to 0.60 to 0.66.

Weights are CC BY-NC 4.0 under FaceForensics++ terms; code is MIT. A score alone cannot establish a video’s authenticity.

02 / Evidence and training

Read the data behind the models

Cards include sources, split details, licenses and limits.

03 / Research collections

Browse by task or language