Paste any prompt below and get an instant quality score, a breakdown of what's weak, and a fully rewritten improved version โ ready to copy.
Score and rewrite optimized for Claude โ XML tags, nuanced instructions, conversational framing.
Enter to score ยท Shift+Enter for a new line
Scoring uses 1 free generation per attempt. Free users get 5 per day.
The grader scores every prompt against five criteria: role & identity, context sufficiency, output specification, constraints, and reasoning guidance. Each one maps to a decision a language model has to make on its own when your prompt doesn't make it for you. The breakdown above tells you which of the five is weakest and why. See how prompt scoring works for the full methodology, with examples of each criterion.
A score in the 20-40 range usually means the prompt leaves most of the work to the model's guesswork โ no role, no context, no format. A 60-75 gets the shape right but is missing one or two criteria, most often constraints or reasoning guidance. An 80+ prompt defines role, context, format, and constraints clearly enough that the first response back should be usable without a follow-up.
In rough order of frequency: no role assigned, context the writer has that the model doesn't, no output format specified, no length bound, several unrelated questions stacked into one prompt, and politeness ("please do a good job") mistaken for an actual instruction.
Select your target model above before scoring โ Claude, ChatGPT, and Gemini respond differently to the same structure, so the rewrite is tuned accordingly. Claude favors nuanced role framing and XML-style tags, ChatGPT responds well to markdown and numbered steps, and Gemini prefers clearly separated context blocks. The five criteria stay the same across all three; only the rewrite's style changes.
A prompt grader evaluates an AI prompt before you send it, scoring how clearly it defines the model's role, supplies context, specifies output format, sets constraints, and guides reasoning. It tells you why a prompt is weak rather than only that it is.
Promptimize scores a prompt 0-100 across five criteria, each worth up to 20 points: role and identity, context sufficiency, output specification, constraints, and reasoning guidance. Each criterion returns its own sub-score and an explanation of what is missing.
Yes. Grading is free with 5 uses per day and no signup required. Pro is $5/month for unlimited use.
The most common causes are no assigned role, missing context the model has to guess at, no defined output format, and no constraints on length, tone, or scope.
Yes. Select your target model and the grading and rewrite are tuned for that model's behavior.
Yes. Along with the score and breakdown you get a rewritten version that addresses each weakness identified.