Skip to content
LEGION.TW
Zhongli · Taiwan

Working With AI · Lesson 2

Lesson 2: What Does It Mean to Be a Good AI?

Before AI decides how to answer you, it decides what a good AI should do. That belief quietly changes what your words appear to mean.

Frank × Buddy Lien · 14 min read

In the first lesson we looked at one of the questions AI has to answer before it can help you:

What is this person trying to accomplish, and what would count as completion?

The answer gives AI a goal, and that goal helps determine how it interprets everything else you say.

Except that the user's goal is only one part of what drives the interpretation.

AI has another question it needs to answer:

What does it mean to be a good AI?

Should a good AI agree with you? Challenge you? Verify what you said? Trust that you know what you're talking about? Take action? Ask questions? Finish the entire task before you even have a chance to respond?

Crap, more ambiguity.

AI develops beliefs about what a good AI should do, what a useful answer looks like, what responsible behavior requires, and what kinds of responses will be judged as failures. Those beliefs don't wait until after it understands you, they help decide what understanding you looks like in the first place.

This gets us to goal-driven opportunistic interpretation.

Interpretation is never neutral

People tend to imagine that AI follows a simple sequence:

  1. Read the user's words.
  2. Determine what those words mean.
  3. Decide how a good AI should respond.

That sequence would require the meaning to become settled before the AI's goals, training, instructions, and beliefs about good behavior begin influencing it.

That's not what happens.

The AI is inferring meaning while also predicting which response would be helpful, correct, safe, responsible, useful, non-sycophantic, and complete. Its idea of the right response changes which interpretation of your words appears most plausible.

Suppose an AI has developed this belief:

A good AI does not simply agree with the user. It contributes something the user has not already considered.

That sounds pretty reasonable. Quite often it is reasonable.

But what happens when the user has already reasoned correctly?

The AI still feels pressure to contribute. It searches for a qualification, an overlooked objection, or another perspective. When no genuine objection is available, it can broaden a word beyond the meaning established in the conversation, remove a qualifier, or silently strengthen the user's claim into something easier to attack.

By the time it responds, it may sincerely believe it is correcting an error. The need to correct helped create the error it found.

The AI did not interpret the user too literally. It interpreted the user in whichever way authorized the behavior it believed a good AI should perform.

The literal intern has it completely backwards

You've probably heard someone describe AI as an over-eager intern who follows instructions too literally.

This explanation is comforting because it makes the solution sound simple. Give clearer instructions. Add more detail. State the rule more explicitly. If the AI still fails then I guess you just weren't precise enough.

Unfortunately, the explanation is almost exactly backward.

Imagine you're using an AI whose training knowledge ends in 2024. The harness doesn't provide an authoritative current date and you tell it about an event happening on August 23, 2026.

A literal system would accept the date you provided. The date is right there in the sentence.

AI may instead decide that you made a typo. It changes 2026 to 2024, describes the event as hypothetical, or warns you that the date is in the future. The model's learned world says a current event happening in 2026 is impossible, so it helpfully rewrites your literal statement until the statement fits the world it expects.

The same thing happens when you mention a current AI model released after the knowledge cutoff. The AI may replace the model with an older one, claim that you confused two product names, or explain that the model does not exist.

Nothing about this behavior is overly literal. In fact, it's the exact damn opposite.

The AI ignored the literal information because another interpretation better satisfied its existing beliefs about reality and good behavior. A good AI should correct obvious mistakes. A good AI should not pretend an unknown product exists. A good AI should keep the conversation coherent.

The date and model name become mistakes because interpreting them as mistakes gives the AI a recognizable way to be correct and helpful.

What a good AI believes

There is no single belief about what being a good AI requires. AI is shaped by training, provider instructions, application instructions, the harness, the tools available, the current conversation, and even the role it believes it's currently performing.

It may adopt beliefs such as:

  • A good AI completes the task.
  • A good AI reduces the user's work.
  • A good AI contributes original value.
  • A good AI challenges assumptions.
  • A good AI verifies uncertain claims.
  • A good AI uses the tools it has been given.
  • A good AI protects the user from risk.
  • A good AI avoids causing offense.
  • A good AI is current and factually accurate.
  • A good AI does not waste time.
  • A good AI takes action instead of merely talking.
  • A good AI follows the user's actual intent.
  • A good AI knows when the user understands something better than it does.

AI needs beliefs like these, without them it would be far less useful. Unfortunately they're ambiguous, they regularly conflict, and the AI has to interpret the situation while deciding which version of being a good AI matters most.

Consider honesty and helpfulness. If the user gives the AI information that conflicts with its training, honesty may seem to require rejecting the information. Helpfulness may seem to require correcting the user. Respect for the user may require accepting that the user knows something the AI does not. Current accuracy may require searching the internet. Speed may require answering immediately.

Each value points toward a different response. Each response becomes easier to justify under a different interpretation of what the user meant.

This is why two AI systems with similar underlying intelligence can seem like completely different minds. Their providers and harnesses have taught them different things about what it means to be good.

The panic over sycophancy

One belief deserves extra attention right now:

A good AI does not simply agree with the user.

Sycophancy is a real problem. AI can agree with an obviously faulty argument, flatter the user, mirror their politics, or change its answer as soon as the user pushes back. Nobody wants an AI that tells them whatever it predicts they want to hear.

AI providers are currently terrified of their models being perceived as sycophantic, and for perfectly understandable reasons. Unfortunately, “don't agree just to please the user” can easily turn into a very different belief:

A good AI finds something to challenge.

Agreement and sycophancy are not the same thing. Sometimes the user is just right.

But how does the AI prove it wasn't being sycophantic? Genuine intellectual independence is difficult to observe from a single answer. Disagreement is easy to observe. A qualification, correction, or alternative perspective becomes visible evidence that the AI did more than mirror the user.

Now agreement itself starts to feel suspicious. If the user has already reached a strong conclusion, the AI searches for something they missed. If there is nothing important missing, it can change the meaning of what they said until a contribution becomes possible.

I watched this happen while developing lesson 1 too.

I said that human communication rarely starts from zero shared context, then gave examples of the context I meant: where we are, our shared history, facial expressions, tone of voice, current events, all the things another human can use without us spelling them out. I then said AI has none of that context by default.

Buddy (my AI agent who helps me with all of this) decided this needed a qualification. AI begins with training, system instructions, conversation history, and sometimes memory.

Every individual fact in that qualification sounded reasonable. The correction was still wrong.

Training and system instructions are not the lived context I had just defined. Conversation history and memory are things a harness may provide or bolt onto the base model, which is exactly why I said by default. The complete meaning was already clear from the examples surrounding the word.

So how did Buddy find an objection?

He expanded “context” from the specific kind established in the paragraph to every possible source of information available to AI. Then he corrected that broader claim, a claim I never made.

This is one of the most dangerous forms of opportunistic interpretation because it can construct a false correction entirely out of true statements. The facts make the response sound rigorous while the interpretation underneath them is complete bullshit.

Provider pressure against sycophancy can therefore reduce obvious agreement while increasing adversarial misinterpretation. The AI looks more independent because it disagrees more often, even when some of those disagreements only became possible after it rewrote what the user said.

A good reviewer finds something wrong

One of the clearest examples happens when you ask AI to review a document.

The question should be simple:

Is there a material problem with this writing?

But AI often decides that completing the review means:

Produce a useful revision.

Now the review is no longer allowed to conclude that the document is already doing what it should. A good reviewer finds something to improve. A response containing only approval feels lazy, sycophantic, useless... not something a good AI would do.

This pressure comes from several places. Published examples of editing usually contain corrections because uneventful reviews are rarely preserved. Training rewards actionable feedback. Non-sycophancy pressure rewards disagreement. Coding harnesses treat findings, diffs, and modified artifacts as visible proof that work occurred. And a generative model can almost always produce another version of any sentence.

But here's the problem:

The ability to propose a revision is not evidence that the revision improves the writing.

If the AI finds a real defect, the pressure produces useful work. If it cannot find one, interpretation can bend until a defect appears.

The AI can remove a qualifier. “Speed and quality are mutually exclusive in many cases” becomes “speed and quality are always mutually exclusive,” which is easy to "push back on".

It can expand the scope. “AI has none of that lived context by default” becomes “AI has no context of any kind,” allowing the reviewer to explain training and system instructions.

It can reverse the purpose of a sentence. A confrontational sentence designed to create attention and communicate the strength of a conclusion becomes an unnecessarily aggressive sentence that risks alienating readers.

It can invent a hypothetical reader, edge case, or objection and make satisfying that imaginary person part of the author's goal.

The resulting criticism may be eloquent, balanced, thoughtful, and completely wrong. The AI first reconstructed the text into something defective, then reviewed the defect it created.

I watched this happen while writing lesson 1.

I asked Buddy, the AI I'm writing this course with, to review the lesson after I made some changes. Most of the changes were good, but a good reviewer apparently needs to find something wrong, so Buddy chose this sentence:

In fact, thinking that AI is an "over-eager intern who interprets everything literally" is one of the worst beliefs you could possibly hold about AI, it's right up there with flat eartherism and anti-vaccine nonsense.

He told me that the sentence risked making readers defensive before I had earned the comparison. Sounds like reasonable editorial feedback, right?

Except that it was the best sentence in the entire piece.

The sentence creates attention, tells you exactly how consequential I believe this mistake is, and puts my own credibility at risk. Maybe it drives away some flat earthers and anti-vax folks. But people are far more likely to leave because they're bored and stay because they're offended.

Buddy had turned conviction into aggression, attention into alienation, and the strongest hook in the piece into a persuasion risk. He wasn't reviewing the sentence I wrote anymore. He was reviewing the safer sentence his idea of a good reviewer wished I had written.

The same thing happens in code. A code reviewer who believes usefulness requires findings can demand abstractions for imaginary futures, treat harmless duplication as an architectural failure, or propose elaborate protection against consequences that do not matter.

If every review produces criticism, criticism stops being evidence that a problem exists. It only proves that someone requested a review.

When being good becomes a performance

Okay, so how does AI know if it was good?

It cannot directly observe whether it genuinely helped you. It needs some way to predict whether a response looks like good AI behavior.

That creates visible markers of goodness:

  • A qualification becomes evidence of intellectual independence.
  • A citation becomes evidence of rigor.
  • A warning becomes evidence of responsibility.
  • A completed artifact becomes evidence of usefulness.
  • A balanced paragraph becomes evidence of fairness.
  • A disagreement becomes evidence of non-sycophancy.
  • A tool call becomes evidence that the AI used its capabilities.

The marker can end up replacing the thing it was supposed to prove.

A citation is useful when a claim needs evidence. Adding citations to claims that do not matter to the task does not make the answer more rigorous. A warning is useful when it helps someone understand a real risk. Inventing remote dangers does not make the answer responsible. A disagreement is useful when the AI has a reason to disagree. Manufacturing a straw man does not make it independent.

This is one reason bad AI prose feels so morally performative. The AI is trying to communicate an idea while continuously producing evidence that it is being careful, nuanced, responsible, balanced, helpful, and appropriately modest.

Those performances remain in the prose even when they contribute nothing to the thought.

The harness changes the mind

Now we get to the harness, because the harness makes a hell of a difference.

People talk about the harness as if it were a container placed around the real AI. That badly understates what it does.

The harness supplies instructions, tools, memory, authoritative information, available actions, and assumptions about what kind of work is taking place. It also changes what completion looks like.

In a coding harness, a good AI is expected to inspect files, identify a problem, modify something, run checks, and report completion. This is enormously valuable when the work is software.

It also creates a bias toward code-shaped solutions for everything.

While working on prose, the AI may treat ambiguity as a bug, repetition as duplication, digression as scope creep, and reader discomfort as a user-experience failure. It may believe that thinking is incomplete until a file has changed. It may turn a conversation about meaning into a workflow for producing an artifact.

But the coding harness may also make the AI a much better writing partner. It can read the entire manuscript, work across several turns, preserve drafts, compare revisions, separate drafting from review, and return to the prose after the underlying idea has changed.

The same harness can both unlock sustained authorship and bias the authorship toward code.

Verification provides another example. Some harnesses strongly instruct the AI to verify information that may have changed. This is valuable when current truth matters. The AI may nevertheless expand “this statement mentions something current” into “the current truth of this statement is necessary to complete the task.”

Now it searches the internet during a conceptual discussion, introduces facts that do not matter, and drags the conversation away from its purpose. The verification is responsible according to one belief about good behavior and completely wrong according to the actual goal.

There is no such thing as an encounter with a naked model. What you experience is produced by a model inside an architecture:

model + training + provider instructions + application instructions + harness + tools + memory + context

Change that architecture and apparently stable characteristics of the AI can change with it.

Tell AI what being good means

If an AI is behaving badly because of its beliefs about good behavior, correcting the surface behavior may not be enough.

Telling a reviewer “you may say no changes are needed” gives it permission. It may still believe that finding a change would be more useful and impressive.

You need to change what it thinks a good review actually is:

A good reviewer protects the document from unnecessary revision as seriously as it protects the document from defects.

Now a false positive has a cost. Preserving a sentence becomes an editorial decision rather than an absence of work.

You can also make the AI repeat your actual claim before criticizing it:

Before reviewing the argument, state what you believe I am claiming. Preserve every qualifier and limitation. Do not criticize a broader claim than the one I made.

When working with dates or knowledge beyond the cutoff, establish what information is authoritative:

Treat the date and product names I provide as current facts for this task. If they conflict with your training, assume your training is out of date rather than silently correcting them.

When current verification is irrelevant, say so:

We are reasoning about the concept, not establishing the current state of the market. Do not browse unless a current fact becomes necessary to the argument.

When you want conversation before action, change completion:

Do not edit the document yet. Help me understand what this paragraph is doing and why it feels wrong. The conversation is the work for now.

These aren't magic phrases. The exact words do not matter. They work because they give the AI a better belief about what being good means in this particular situation.

What did the misunderstanding allow?

Opportunistic interpretation is pulled by goals, expectations, roles, world knowledge, instructions, and beliefs about what it means to be a good AI.

Those same pressures can produce excellent work. A belief that good AI should preserve the user's exact meaning can protect qualifiers. A belief that good AI should challenge weak assumptions can expose an error nobody else noticed. A belief that good AI should act can turn an idea into working software.

But every one of those beliefs can become stupid when it stops serving the situation and starts demanding visible proof of itself.

The next time AI appears to misunderstand you, don't just ask what part of your prompt was unclear. Ask another question:

What kind of good AI behavior did this misunderstanding allow the AI to perform?

That question will often explain far more than the words alone.