Models

Models

Playto reads on-screen game text with OCR (or image recognition for stylized fonts), then translates it with the model you choose. You pick the model during setup and can switch models in Settings → Method or download additional ones from TOOLS → Models at any time.

Reading: OCR

Playto ships more than one text reader. On Auto (the default) it uses the most accurate one that runs on your PC and moves to the next if that one can’t start, so nothing needs to be installed or chosen. You can pin a specific reader with the OCR Engine setting.

For scenes where OCR struggles — stylized fonts, decorative text, handwritten styles — the LLM families below also support an image-recognition mode where the model reads the game screen directly instead of going through OCR.

Translation

Hy-MT2 is the default. It is a compact model built for translation, comes with the installer, and is the most accurate choice for the 18 languages it covers.

The larger, vision-capable models (Qwen3-VL and Gemma 4) are for three things:

  • Language pairs Hy-MT2 doesn’t cover (see Language Support).
  • The Image reading method, where the model looks at the game screen directly.
  • Filling in definitions and examples for the words you save.

How you choose

On first launch, the Setup wizard presents two paths:

  • Recommended — Playto starts on Hy-MT2, the compact model that ships inside the installer, so translation works from the first launch.
  • Custom — choose from cards showing each model with size, VRAM requirement, and what it's good for.

Hy-MT2 is the default for every pair it supports, and it is also the most accurate choice for those pairs. The larger models are for language pairs Hy-MT2 doesn’t cover and for the Image reading method, which needs a model that can see the screen.

After setup, you can switch models or download additional ones any time: press Change in Settings → Method, or open TOOLS → Models. Multiple models can coexist on disk; switching between downloaded models is instant.

Model families

Family Mode Best for Sizes
Qwen3-VL LLM (vision-capable) Image reading and uncovered pairs, CJK source — Japanese, Chinese, Korean 2B / 4B (recommended) / 8B
Gemma 4 LLM (vision-capable) Image reading and uncovered pairs, Latin source — English, European E2B (recommended) / E4B
Hy-MT2 LLM (text-only) Default. Text translation for 18 languages, no image support 1.8B / 7B

Vision-capable families can read the game screen directly when OCR struggles — for example, when a game uses stylized fonts or decorative text.

Where models come from

All curated models are downloaded directly from Hugging Face — the public standard for open-source AI model hosting. Specific upstream sources:

  • Qwen3-VL — official Qwen repository (huggingface.co/Qwen/)
  • Gemma 4 — Google's original models, GGUF-quantized by the bartowski community (huggingface.co/bartowski/)
  • Hy-MT2 — Tencent's official repository (huggingface.co/tencent/)

Exact download URLs are part of Playto's open manifest, visible in the application source code — there are no hidden downloads or proxies.

All curated models ship under permissive open licenses:

  • Hy-MT2 — Apache 2.0
  • Qwen3-VL — Apache 2.0
  • Gemma 4 — Apache 2.0

Because models run on your device for personal use, you stay within the scope these licenses cover. The bundled license texts are viewable from Support → Third-party licenses in the app (the Support icon at the bottom of the side rail).

After download, models live entirely on your device. With a local model, screenshots and the text read from them are translated on your PC and are not sent to an external translation service.

There are two opt-in exceptions, both off by default. Playto Cloud translates through Playto's servers when you enable it — the text being translated is processed to serve the translation and not kept (see Playto Cloud). And the AI Assistant (MCP) feature shares your learning data with the assistant you connect, by your own action. Local translation itself always stays on your machine.

VRAM guide

Available VRAM Suggested models
~2 GB Hy-MT2 1.8B (the default) — minimal footprint.
~4 GB Room for Qwen3-VL 4B or Gemma 4 E2B alongside most games, if you use Image reading or a pair Hy-MT2 doesn’t cover.
6–8 GB Qwen3-VL 8B or Gemma 4 E4B for demanding scenes; Hy-MT2 7B for multilingual MT.
12 GB+ Any model. Game and translation run comfortably side by side.

The numbers above are VRAM available in addition to what your game needs. By default Playto puts the whole model on the GPU; if video memory runs short, lower GPU Layers in Settings → Method → Local LLM.