How to Score Sales Calls with Jev
Jev by TypeSafe AI brings a different approach to understanding your sales conversations: it reads what was said and returns defined answers your team can use across calls. Buyer objections become categories, discovery quality becomes a score, and qualification evidence becomes a set of fields you can see in your pipeline reviews, including findings from the conversations your managers never have time to open.
What is Jev?
TypeSafe describes Jev as a System One decision model, built to assess language and return an answer from the options you’ve defined. Its output is a category, score or probability rather than a written response, which gives it a different role from a conventional text-generating LLM. That design suits classification: “call me next quarter” can become a timing objection, while “please send pricing” gives your team a different kind of follow-up.
Jev can assess several independent questions against the same conversation in parallel, so your buyer’s problem, budget status and next-step commitment can each have their own result. Probability estimates also give the evaluation a way to express uncertainty when the evidence fits more than one answer, allowing your team to distinguish clear findings from those needing a closer look.

Why Jev matters for sales teams
Your managers already classify conversations as they listen, recognising commercial objections and noticing when a buyer hasn’t committed to a next step. Jev makes those repeated assessments worth exploring across a larger share of your calls, so your managers can spend more of their review time on the findings that need attention.
A timing objection could guide your rep’s next conversation while helping your marketing team understand why campaign replies aren’t progressing. Across your pipeline, approval gaps could show where discovery is repeatedly falling short, while recurring product questions could tell enablement where your reps need better support. The same observation becomes useful to several teams because they can work from a consistent field.
At TypeSafe’s listed rate of $0.042 per million input tokens on 19 September 2026, 10,000 calls using 12,000 input tokens each would cost $5.04 in Jev charges. Tokens are the text units used for model billing, covering your transcript and questions, and the low model cost creates room to assess more conversations and more dimensions of each call.
Sales Scorecards
Your scorecard should reveal something you can coach or act on, which is why separate categories and scores are more useful than one judgement on whether a call went well. Your rep may have uncovered a serious problem and secured another meeting while also establishing that your buyer has no approved funding; those findings belong in different conversations about the opportunity.
If your buyer says “three people spend six hours a week reconciling invoices, but finance hasn’t approved funding”, a scorecard could represent that as three out of three for problem specificity and a budget status of “not approved”. Your buyer has quantified a problem worth exploring, while your rep still has work to do on the purchase process.
Jev’s Score questions assess evidence against written levels, giving a high score the meaning your sales team assigns to it. General dissatisfaction might sit at the bottom of your scale, a specific operational problem above it, and a problem with a clear consequence and measured impact at the top. Choice questions handle categories such as approval status, while Noul returns the estimated probability that a statement is true, such as your buyer explicitly accepting a next step.
Together, those results help you distinguish rep performance from opportunity quality: a well-handled call may uncover a poor-fit deal, while a strong opportunity can still expose gaps in your rep’s discovery. Our approach to consistent evaluation questions and fields gives those distinctions a repeatable meaning across your calls and coaching conversations.

The quality and correctness of your answers
Your rep answering a question establishes that the topic was covered, but you still need to know how well the answer addressed your buyer’s concern and whether it was correct. Suppose your rep says single sign-on is included in Starter, while your product guide says it’s available only on Enterprise: your scorecard can recognise the answer and separately flag the product claim as incorrect.
If your buyer asked because their IT team needs to review access, a brief “yes, it’s included” leaves the concern largely unexplored, whereas a useful response would explain the capability and establish what their IT team needs to assess. Jev can evaluate answer quality separately from correctness when your conversation and the relevant guidance are supplied together.
Your product guide, pricing rules and playbooks give these judgements their business meaning, because the same answer can be accurate for one plan or customer segment and misleading for another. Our article on grounding call scoring in product, pricing and ICP knowledge explains why that context changes what your scorecard should reward.

What your calls leave unanswered
Your buyer saying IT needs to review the purchase can reveal a useful discovery gap if your playbook expects an identified review owner and your calls haven’t established one. Jev can assess whether the review is required and whether its owner has been named, leaving your manager with a specific finding to discuss with your rep. Missing ownership remains different from a confirmed blocker, just as budget never coming up is different from your buyer saying funding hasn’t been approved.
Across your pipeline, those distinctions give you a view of discovery coverage that an average call score would hide. Your managers can see which questions keep going unanswered and connect that pattern to opportunity outcomes, while checking earlier conversations before treating a gap as something your team has never covered. That’s the value of consistent conversation fields for BI: the evidence supports your reporting as well as your next coaching conversation.

Where deeper reasoning helps
Jev’s focused decisions cover a valuable part of your scorecard, while some sales judgements require connecting several facts and weighing competing explanations. Checking whether a feature is included in a plan is a direct comparison; deciding whether a different plan is the right recommendation may involve your buyer’s requirements, budget constraints and an exception in your pricing policy.
TypeSafe documents weaknesses in Jev 1.13 around questions requiring several linked reasoning steps, indirect wording and long inputs containing irrelevant material, and warns that related answers aren’t guaranteed to agree with each other. A frontier reasoning model may add value where your buyer’s position changes across calls or several documents need interpreting together, and it can also turn assessed findings into a written coaching explanation.
The benefit of that extra reasoning depends on how the combined assessment performs against conversations your managers have reviewed, including the difficult cases. Your comparison needs the full cost of model calls and review work alongside the quality of the decisions, because neither a neatly formatted result nor a confident explanation establishes that your evidence has been interpreted correctly.
Why knowledge bases are essential plumbing
Your rep’s answer is correct only in relation to what your business actually offers, and your model needs access to that information when it makes the assessment. Product capabilities, commercial terms and approved exceptions all affect the judgement, along with the expectations your playbook sets for that buyer and stage of the sale.
The conversation still follows the same path from transcript to model to structured data, with your knowledge base connected alongside the assessment. Your evaluation draws in the relevant passages from your documents, giving the model company context to compare with what your rep said. That connection is what Knowledge grounding provides.
For a question about single sign-on, the relevant passage might come from your product guide; for an onboarding promise, it might be the part of your commercial terms that distinguishes included support from paid migration work. Having those references available lets the assessment identify a specific discrepancy, rather than judge the answer against a general idea of what a software company normally provides.
Your enterprise and self-service offers may have different terms, your pricing may vary by region, and an older call may have taken place before a product change. A model can apply the wrong document faithfully and still produce an unfair score, so the source material needs to match the circumstances of your conversation. Where your documents don’t establish the answer, correctness should remain unresolved.
Finding the relevant material, keeping it current and connecting it to the right evaluation are ongoing responsibilities. As your scorecards spread across teams, those document connections become part of the system your sales data depends on, with changes to your guidance affecting the meaning of the results. A more capable model still needs that plumbing in place.

How Semarize brings this together
We built Semarize to make your company-specific evaluations reusable across teams, with the scoring questions, supporting knowledge and versions managed together. Bricks hold individual questions, while Kits group them around discovery quality, objection handling or another use case, with the relevant knowledge bases attached. Your managers can work from shared definitions of quality and correctness, and your operations team can maintain those definitions centrally.
The Semarize API brings that framework into your existing tools: your workflow sends a conversation and selects a Kit, then Runs apply its checks and return consistent fields with reasons, confidence and evidence. Those results can support your CRM, reporting and coaching processes while your team continues working in the tools it already uses. Versions let you test changes, choose when integrations adopt them and identify which assessment produced an earlier result.
Semarize provides a managed route to conversation evaluation, separate from a direct Jev implementation. For your sales organisation, the benefit is having company-specific judgements available as usable data, with a practical way to maintain the scorecards and document connections behind them as your team and business change.
Common questions
What makes Jev different from a conventional LLM?
TypeSafe describes Jev as a decision model that understands language and returns defined categories, scores and probabilities rather than generating prose. That suits repeated judgements across your conversations, such as classifying an objection or assessing whether your buyer confirmed a next step. Written explanations and questions involving more extensive reasoning can be handled separately where they add value.
Can Jev assess more than whether something was mentioned?
Jev's Score questions can assess your conversation against descriptive quality levels, while other questions can compare a claim with supplied company guidance or identify expected evidence that is missing. The usefulness of those assessments depends on clear criteria and the right context, such as your product guide for correctness or your discovery playbook for coverage.
Does a high call score mean your deal will close?
A call score reflects the criteria being assessed, so a clear business problem and an accepted next step can coexist with unresolved funding or approval questions. Keeping those findings separate gives your pipeline reviews more useful evidence. Estimating a chance of closing requires a separate assessment against historical outcomes using information available at the time.
Why does your scorecard need company knowledge?
Your company documents establish which product claims are correct, what pricing rules apply and which discovery points matter to your sales process. In Semarize, Knowledge grounding attaches that context to Kits, allowing the evaluation to assess quality, correctness and missing evidence against your guidance, with unresolved results where your documents or conversations don't provide enough information.
Continue reading
Read more from Semarize
AI Call Scoring Is Theatre Without a Knowledge Layer
AI call scoring that runs on a good LLM with a well-written rubric can look accurate until you test it against what actually happened. The failure isn't one missing check. Every commercial dimension worth assessing has multiple facets, and each facet requires its own grounded document to evaluate properly. A knowledge layer is what makes scoring checkable across all of them rather than plausible about none of them.
Bricks and Kits: the mechanism for stable conversation evaluation
Freeform prompts produce inconsistent evaluation results - scores drift, output shapes change, and you can't tell whether coaching improved anything or whether the rubric moved. Bricks define a locked evaluation schema: one question, one output type. Kits group them into reusable evaluation workflows. The result is schema-stable conversation analysis you control.
Conversation Data Warehouse: Consistent Call Fields for BI
BI teams can't query transcripts. They can't join AI summaries to CRM objects. To make conversation data useful for analytics, it needs to arrive as consistent typed fields - booleans, scores, text fields, lists - with join keys that connect calls to opportunities, accounts, and contacts. This is the pipeline, the schema, and the governance model that makes sales call analytics possible in your warehouse.