The cost of an AI rejection

In LIFT's documented Upwork workflow, writing the proposal is step six. First come the brief, an opportunity score, a commercial assessment, a strategy and the choice of relevant work to show.

That order matters when AI becomes cheaper. The first useful question is whether it can help choose the work worth pursuing. A polished proposal never gets written for a brief the system has already dismissed.

OpenAI's 22 September launch of GPT-6 Sol and Luna gives that question fresh relevance. Published standard API prices for prompts up to 272,000 input tokens put Luna at $0.10 per million input tokens and $0.50 per million output tokens, compared with Sol at $2 and $10 respectively. Those are US-dollar token rates, not the cost of a completed commercial assessment. They do not establish which model will judge a studio's opportunities well. [1]

This article applies that development to LIFT's recorded acquisition rules. It is sourced analysis, not a report of testing these models. We have no measured conversion improvement, saving or missed opportunity to claim.

A budget filter has exceptions

LIFT's Riker instructions contain something more useful than a universal minimum budget: exceptions that require attention.

They normally exclude fixed-price work below $750, but allow review of a simple existing Squarespace refresh or a credible repeat-client opportunity. They also say a $1,200 project from a client with a substantial spending history on Upwork may be preferable to a $2,500 project from an unknown client. These are examples in our operating instructions, not accounts of actual clients or evidence that the lower-priced job will be better. [2]

The reasoning is commercial. A headline budget does not tell us how much work sits behind it. An existing site with a contained change and a large rebuild can have very different demands. Client history can provide relevant evidence, while an unfamiliar client may simply need more investigation.

A screening model that applies the budget threshold but loses the exception has misunderstood the task. Its recommendation might still look admirably clear: a score, a red flag and a confident instruction to skip.

That clarity is dangerous if it conceals the part of the brief that deserved consideration.

Unknown does not mean unsuitable

Screening creates two kinds of error. A poor-fit brief can be recommended, consuming time that could have gone elsewhere. A potentially suitable brief can be rejected, disappearing before anyone considers it properly.

The second error is harder to notice. A rejected opportunity does not return later to explain that its scope was manageable or its client worth talking to. Equally, its advertised budget cannot be counted as lost revenue. LIFT might never have won it, and the work might not have been profitable.

That uncertainty is why 'clarify' needs to remain a real outcome.

A missing budget is a missing budget. It is not proof that a client cannot pay. A short brief may lack enough information to estimate delivery effort. An unfamiliar platform may be a firm requirement, or the client may be asking for a recommendation. Those differences should be established from the supplied evidence, not completed by a model that wants every field filled.

Riker's rules also ask for likely hours and effective hourly value. That is a useful question, but an estimate built on missing scope should be labelled provisional. A confident number does not make its assumptions sound.

The useful output may therefore be a precise unresolved question: does this existing site need a contained refresh, or does the brief imply a rebuild? That question can change the decision more than another paragraph of sales copy.

Test the rejected pile

Rob's recorded decision on 27 June was to begin with manually supplied briefs, a scoring sheet, client-quality checks and proposal-outcome tracking. He ruled out automated submissions and messaging. That record establishes a starting design, not evidence that it has improved results. [3]

It also suggests a practical way to assess a cheaper model before giving it influence over the shortlist.

Take a small collection of real briefs the studio is permitted to use. Have the person who makes the commercial decision classify each independently as worth reviewing, unsuitable or needing clarification, and record why. Then give the same material and the same written rules to the candidate model.

Do not evaluate only the briefs it recommends. Open the rejected pile too.

For each disagreement, check what happened. Did the model miss an explicit exception? Treat absent client history as adverse history? Assume a platform was negotiable when the brief made it compulsory? Underestimate the work behind a reasonable-looking fixed fee? Or identify a genuine concern the human overlooked?

The human's first answer is not automatically correct. The point of comparison is to expose the reasoning and return to the brief. Where reasonable people would disagree, preserve the disagreement instead of forcing a score to imply certainty.

This is a proposed test. LIFT has not completed it for these releases, and it would not establish a universal ranking of models.

Reviewing everything has a cost too

There is a fair objection: if somebody must inspect every rejection, where is the benefit?

During evaluation, that inspection is how the studio learns which decisions it can trust. It need not mean reading every rejected brief forever. But reducing review without that evidence risks hiding mistakes rather than reducing them.

The result may justify a smaller assignment. A lower-cost model could extract explicit budget, platform and deliverables while leaving ambiguous commercial recommendations to a person. It could highlight the wording that supports an exception without deciding that the exception guarantees a good opportunity.

If the model reliably handles clear cases but struggles with incomplete scope, that is a useful finding. Route incomplete scope for review. If it cannot distinguish unknown information from a disqualifying fact, keep it out of rejection decisions until the problem is addressed.

This is more specific than asking whether its writing sounds good. A beautifully phrased explanation of the wrong rejection is still the wrong rejection.

Keep the reason beside the decision

For this use, the assessment record should retain the original brief, the rule applied, the supporting words, the recommendation and any unresolved question. Record the model and version too, so a later change can be checked against the same material.

Where a proposal proceeds, later interviews, wins, losses and actual delivery effort can inform future judgement. Where it does not, avoid inventing a counterfactual success rate. The record can show whether the decision followed the evidence available at the time; it cannot reveal the contract that might have happened.

Cheaper inference creates room to examine briefs more carefully. Whether it should also decide what Rob never sees is a separate question.

LIFT's own rules make the test concrete: preserve the exception, distinguish uncertainty from poor fit, and show the reason before closing the opportunity. The value of screening is in the work it helps a studio choose. Some of that value is hidden in the pile marked 'no'.

Sources and editorial basis

OpenAI: Introducing GPT-6 Sol and Luna and API changelog.

[1] OpenAI, Introducing GPT-6 Sol and Luna, and the API changelog, 22 September 2026. Provider-reported prices; no LIFT performance claim.

[2] LIFT's Riker training notes and Lift Upwork Agent workflow, reviewed 25 September 2026. Internal source artefacts, not measured outcomes. The budget figures describe recorded rules and are not a current price quotation.

[3] LIFT CORP Founding Board Answers, 27 June 2026. Dated record of Rob's manual-input, scoring and tracking decision.

Discuss your website or AI workflow with LIFT Brandworks.

Lift BrandWorks

Hi, I’m Rob, a Squarespace Gold Circle Partner and designer with a focus on fast, high-converting websites for growing businesses. I specialise in clean design, persuasive copy, and smart user flows that turn traffic into trade.

I’ve worked across hospitality, wellness, coaching, and creative sectors, delivering custom builds without templates, every site is tailored, mobile-optimised, SEO-ready, and built for growth. I’m also a trained coach and great communicator, making the whole process easy and collaborative from start to finish.

Need help launching or refreshing your site? I offer full design, development, and optional chatbot integration to help your site work harder for you.

https://www.liftbrandworks.com
Previous
Previous

AI should take work off the table, not follow you home

Next
Next

AI slop is becoming a distribution risk