Claude Code Guide
Set up Flaq AI Claude models and explore Claude Code skills
Free to try Grok 4.5 Image-to-Text API powered by X-AI for visual Q&A, OCR, image analysis, and multimodal reasoning through Flaq AI. Explore visual workflows and turn images into actionable insights. Lower cost for OCR than Claude Fable5.
const response = await fetch('https://api.flaq.ai/api/v1/chat/completions', {
method: 'POST',
headers: {
Authorization: 'Bearer YOUR_API_KEY',
Accept: 'text/event-stream',
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: 'grok-4.5-image-to-text',
messages: [
{
role: 'user',
content: [
{ type: 'text', text: 'Describe the image and extract any visible text.' },
{
type: 'image_url',
image_url: {
url: 'https://example.com/sample-image.jpg'
}
}
]
}
],
stream: true,
max_tokens: 2048
})
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
let assistantText = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const frames = buffer.split('\n\n');
buffer = frames.pop() || '';
for (const frame of frames) {
const lines = frame.split('\n').filter(Boolean);
let eventName = 'message';
const dataLines = [];
for (const line of lines) {
if (line.startsWith('event:')) {
eventName = line.slice(6).trim();
} else if (line.startsWith('data:')) {
dataLines.push(line.replace(/^data:\s*/, ''));
}
}
const raw = dataLines.join('\n').trim();
if (raw === '[DONE]') {
console.log('\nFinal text:', assistantText);
continue;
}
let payload;
try {
payload = JSON.parse(raw);
} catch {
continue;
}
if (eventName === 'error' || payload.error) {
const msg = payload.error?.message ?? payload.message ?? 'Chat request failed';
throw new Error(msg);
}
const delta = payload.choices?.[0]?.delta;
if (delta?.content) {
assistantText += delta.content;
console.log(assistantText);
}
}
}
import json
import requests
response = requests.post(
'https://api.flaq.ai/api/v1/chat/completions',
headers={
'Authorization': 'Bearer YOUR_API_KEY',
'Accept': 'text/event-stream',
'Content-Type': 'application/json',
},
json={
'model': 'grok-4.5-image-to-text',
'messages': [
{
'role': 'user',
'content': [
{'type': 'text', 'text': 'Describe the image and extract any visible text.'},
{
'type': 'image_url',
'image_url': {
'url': 'https://example.com/sample-image.jpg'
}
},
],
}
],
'stream': True,
'max_tokens': 2048,
},
stream=True,
)
response.raise_for_status()
event_name = 'message'
assistant_text = ''
for raw_line in response.iter_lines(decode_unicode=True):
if not raw_line:
event_name = 'message'
continue
if raw_line.startswith('event:'):
event_name = raw_line.replace('event:', '', 1).strip()
continue
if raw_line.startswith('data:'):
raw_data = raw_line.replace('data:', '', 1).strip()
if raw_data == '[DONE]':
print('\nFinal text:', assistant_text)
continue
payload = json.loads(raw_data)
if event_name == 'error' or payload.get('error'):
error = payload.get('error') or payload
raise RuntimeError(error.get('message', 'Chat request failed'))
choices = payload.get('choices') or []
if choices:
delta = choices[0].get('delta') or {}
content = delta.get('content')
if content:
assistant_text += content
print(content, end='', flush=True)
curl -N -X POST "https://api.flaq.ai/api/v1/chat/completions" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Accept: text/event-stream" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.5-image-to-text",
"messages": [
{
"role": "user",
"content": [
{ "type": "text", "text": "Describe the image and extract any visible text." },
{
"type": "image_url",
"image_url": {
"url": "https://example.com/sample-image.jpg"
}
}
]
}
],
"stream": true,
"max_tokens": 2048
}'
| Parameters | Price | Original Price | Discount |
|---|
Grok 4.5 Image-to-Text API provides a focused way to use xAI's multimodal Grok 4.5 model for image understanding through Flaq AI. Submit a supported image with an optional text instruction and receive a text response based on the visual context. The current Flaq AI route is designed for single-image analysis, visual question answering, and image-grounded conversation without implying support for general file attachments or image generation.
Note Image interpretations and visible-text extraction can be incomplete or inaccurate, particularly with small text, ambiguous scenes, low-quality images, or specialized content. Review important results before use. This route supports one image and does not expose general file input or image generation.
Grok 4.5 Image-to-Text vs. Grok 4.5 Text-to-Text Image-to-Text accepts visual context and returns text, while Text-to-Text is the more direct choice for requests that contain no image.
Grok 4.5 Image-to-Text vs. Dedicated OCR Dedicated OCR systems are purpose-built for deterministic text extraction and document pipelines. Grok 4.5 can discuss visible content and text in context, but it should not be treated as a guaranteed replacement for accuracy-critical OCR.
Grok 4.5 Image-to-Text vs. Image Captioning Models Captioning models often target short descriptions. A multimodal chat workflow can also respond to a user-supplied question, though output quality remains dependent on the image and instruction.
Grok 4.5 Image-to-Text vs. Other Multimodal LLM APIs Leading multimodal APIs differ in input limits, model behavior, latency, and platform features. Evaluate them with representative images and task-specific acceptance criteria before selecting a provider.
Grok 4.5 Image-to-Text vs. Self-Hosted Vision Models Self-hosted vision models provide deployment control but add infrastructure and maintenance work. Flaq AI offers a managed, single-image chat route for teams that prefer hosted Grok 4.5 access.
Set up Flaq AI Claude models and explore Claude Code skills

Set up Flaq AI GPT models and explore Codex skills
Use Flaq AI LLM models in Hermes Agent
Use GLM 5.2, Kimi K3, and DeepSeek v4 in ZCode
Run DeepSeek Harness with Flaq AI DeepSeek models
Connect AI agents to Flaq image and video generation tools