Translation
Translation
Playto offers multiple ways to translate game text — from classic OCR to AI vision, with local and cloud options.
Text
The lightest mode. Playto captures a region of your screen and reads the text with a built-in text reader, then sends it to a local LLM for translation. Playto ships more than one reader and uses the first one that runs on your PC, so you never have to pick.
Pipeline: Screen capture → OCR text recognition → text stability check → LLM translation → overlay display.
Text stability: Playto waits briefly for the text to settle before translating, and skips re-translation when the text barely changes to avoid flicker.
Text+
The default mode for most games. Text+ runs the same fast OCR pipeline as Text, and adds an image-recognition rescue pass on scenes OCR cannot read — decorative fonts, text drawn into artwork — so those screens get read too instead of staying silent.
Best for: Most games. You keep Text's speed on ordinary dialogue and gain coverage on the screens plain OCR misses.
Image
In this optional mode, Playto sends the screen image directly to a Vision Language Model, which reads and translates it in one pass. It's the heaviest of the three modes — best kept for the cases the OCR-based modes struggle with.
Best for: Stylized fonts, text on complex backgrounds, handwritten-style text, and UI elements that plain OCR struggles with.
Streaming: the overlay updates as the model generates translation tokens, so text appears without waiting for the full response.
Playto Cloud
An optional route that translates through Playto's servers instead of your GPU — no model download, no API key, and a free trial included. It is strictly opt-in: until you enable it, translation runs entirely on your PC.
Best for: starting before the local model download finishes, hardware that can't run local models, and networks where multi-gigabyte downloads won't complete.
Full details: Playto Cloud.
Prompt Patterns
Playto offers three prompt patterns that control how captured text is sent to the AI for translation:
- Auto — Playto decides the best pattern based on the text structure. Best for most games.
- Dialogue — Treats the captured text as a continuous subtitle. Lines are merged into a single paragraph before translation. Best for dialogue-heavy games and cutscenes.
- UI — Each line is translated independently. Best for menus, item lists, and HUD text where lines are unrelated.
Configure per cursor preset or fixed region in the Capture Area card. You can also set per-game overrides in Game Profiles.
Conversation Context
Playto passes the previous translation as context to keep tone consistent across rapid captures.
Local LLM Setup
Playto runs a llama.cpp server locally for translation inference. Models are downloaded with one click and tuned automatically for your GPU — you can change the Backend and GPU Layers yourself in Settings → Method → Local LLM.
Text Reading and Image Recognition can run on separate servers simultaneously, so you can use both modes without switching models.