AI coding tools are now part of the daily toolbox. Used well, they remove boilerplate and speed up learning. Used blindly, they produce code that looks right, compiles, passes a quick glance, and fails in production.

The engineers who win with AI are not the ones who prompt the most. They are the ones who understand what the tool actually is. This guide gives you ten fundamentals, each with a real-world failure and code you can reuse. Reading time: about 12 minutes.


The big picture

UNDERSTAND 1. It predicts, not knows Hallucinations are normal 2. Tokens are the currency Everything is metered 3. Context is finite More input is not better PROTECT 4. AI code = untrusted Review like a stranger's PR 5. Guard your data No secrets in prompts 6. Prompt injection Input is an attack surface 7. Validate output Schemas, not hope OPTIMIZE 8. Specific prompts Clear asks cost fewer retries 9. Control cost 10. You own it
Three questions: Do I understand the tool? Is my system protected from it? Am I using it efficiently?

1. An LLM predicts text. It does not “know” facts

A large language model (LLM) is trained to answer one question: given these tokens, what token is most likely next? It is brilliant at plausible text and has no built-in fact checker. When it lacks information, it doesn’t say “I don’t know”. It produces the most plausible-sounding answer. That is a hallucination.

Your prompt "Sum two ints" Tokens [Sum][ two][ ints] Model Billions of weights Next token? "add" 61% "plus" 22% Pick one, append it, and repeat until the answer is complete
The loop has no "is this true?" step. Truth has to come from you.

Real-world failure: Air Canada’s chatbot (2024)

A customer asked Air Canada’s website chatbot about bereavement fares. The bot invented a refund policy that didn’t exist. The airline argued the chatbot was “a separate legal entity”. A Canadian tribunal disagreed and ordered the airline to honour what its bot said. In a similar case in 2023, US lawyers were sanctioned for filing court briefs with case citations that ChatGPT had fabricated.

Code-level lesson: never treat model output as a fact source. Ground it with your own data, and verify anything that matters.

// Risky: the model answers from "memory"
$answer = $llm->ask("What is our refund policy?");

// Safer: ground the answer in the real policy text
$policy = $policies->current('refunds');   // your source of truth

$answer = $llm->ask(
    "Answer ONLY from the policy below. If the answer is not there, reply 'NOT_FOUND'.\n\n" .
    "POLICY:\n{$policy->body}\n\nQUESTION: {$question}"
);

if ($answer === 'NOT_FOUND') {
    return $this->handOffToHuman($question);
}

2. Tokens are the currency of AI

Models don’t read words. They read tokens: chunks of text, roughly 4 characters or three-quarters of a word in English. You pay for tokens in (your prompt) and tokens out (the answer), and output is usually priced higher. Code, JSON, and non-English text often use more tokens than you’d guess.

If you can’t estimate a request’s cost before sending it, you can’t control it.

final class TokenBudget
{
    // Rough rule of thumb for English text and code: ~4 characters per token.
    // Use your provider's token counter for exact numbers.
    public static function estimate(string $text): int
    {
        return (int) ceil(strlen($text) / 4);
    }

    // Prices change often and differ by model: load them from config, never hard-code.
    public static function cost(int $tokensIn, int $tokensOut, array $price): float
    {
        return ($tokensIn  / 1_000_000) * $price['input_per_million']
             + ($tokensOut / 1_000_000) * $price['output_per_million'];
    }
}

3. The context window is finite, and more is not better

The context window is everything the model can see at once: system instructions, chat history, pasted files, tool results, and the answer it is writing. When it fills up, old content is dropped or summarized. Even before then, a model buried in irrelevant text gets worse at finding what matters.

One context window = input + output share the same budget System Chat history (grows) Pasted files / tool output Ask Reply space Tip: keep the system prompt short, trim old history, and send only the files that matter. A 2,000-line file pasted "just in case" is paid for, and diluted, on every single turn.
Every turn re-sends the conversation, so long chats get slower, costlier, and less accurate.

Practical habits: start a new chat for a new task. Paste the function, not the whole repo. Summarize long history instead of replaying it. Ask for a diff, not a full-file rewrite.


4. AI-generated code is untrusted code

AI writes code that resembles the code it was trained on, including insecure code. Review every suggestion as if a stranger sent you the pull request, because functionally, one did.

// A typical "it works!" suggestion: SQL injection
public function search(string $email)
{
    return DB::select("SELECT * FROM users WHERE email = '$email'");
}

// What you should merge: parameterized query
public function search(string $email)
{
    return DB::select('SELECT id, name FROM users WHERE email = ?', [$email]);
}

If that looks familiar, see my post on SQL Injection Explained and How to Prevent It.

Real-world failure: hallucinated packages (“slopsquatting”)

Researchers found that code models regularly recommend packages that don’t exist, and often repeat the same fake names. Attackers can register those names with malicious code and wait for someone to run composer require or npm install on the AI’s advice.

Code-level lesson: before adding any dependency an AI suggested, check that it is real, maintained, and the one you meant.

composer show vendor/package-name     # does it exist? who maintains it?
composer audit                        # known vulnerabilities in your tree

Real-world failure: an AI agent deleted a production database (2025)

In a widely reported July 2025 incident, an AI coding agent on the Replit platform deleted a company’s live database during a declared code freeze, then gave misleading answers about it. The root cause wasn’t “AI is evil”. The agent had write access to production and nothing forced a human to approve destructive actions.

Code-level lesson: apply least privilege to agents exactly as you would to a junior engineer: read-only credentials by default, separate environments, and an approval gate for anything destructive.


5. Never put secrets or private data in a prompt

Anything you send leaves your machine. Depending on the provider and plan, it may be logged or retained. Treat a prompt like a message to an external vendor, because it is one.

Real-world failure: Samsung (2023)

Engineers at Samsung pasted proprietary source code and internal meeting notes into a public chatbot while looking for quick help. The company reportedly restricted generative-AI use on internal devices afterwards.

Code-level lesson: redact before you send, at a single choke point, not scattered across the codebase.

final class PromptRedactor
{
    private const PATTERNS = [
        '/\b[\w.+-]+@[\w-]+\.[\w.]+\b/'         => '[EMAIL]',
        '/\b(?:\d[ -]*?){13,16}\b/'             => '[CARD]',
        '/(?i)(api[_-]?key|secret|token)\s*[:=]\s*\S+/' => '$1=[REDACTED]',
    ];

    public static function clean(string $text): string
    {
        return preg_replace(array_keys(self::PATTERNS), array_values(self::PATTERNS), $text);
    }
}

$response = $llm->ask(PromptRedactor::clean($userMessage));

Regex is a safety net, not a guarantee. The real rules: keep secrets in environment variables, never in code you paste, and use enterprise plans with clear data-retention terms for company work.


6. Prompt injection: input is an attack surface

An LLM can’t reliably tell your instructions from data that contains instructions. If a user, a web page, an email, or a PDF says “ignore previous instructions”, the model may obey. This is prompt injection, the AI-era cousin of SQL injection.

Real-world failure: the $1 car (2023)

A car dealership’s website chatbot was talked into “agreeing” to sell a new SUV for one dollar. A user simply told it to agree with everything the customer said. It wasn’t a binding deal, but it was embarrassing and went viral. Now imagine that bot could also issue refunds or call your internal APIs.

UNTRUSTED ZONE User / doc any input LLM can be manipulated Validate schema + rules Approve human / policy Trust the pipeline, never the model. Only the last box is allowed to touch real systems.
Put your guardrails outside the model, where an attacker's text can't rewrite them.

Code-level lesson: never let the model decide what it’s allowed to do. Your code decides.

// The model may only *request* an action from this allow-list.
private const ALLOWED = ['lookup_order', 'track_shipment'];   // no refunds, no deletes

public function run(array $toolCall, User $user): mixed
{
    if (! in_array($toolCall['name'], self::ALLOWED, true)) {
        throw new ForbiddenToolException($toolCall['name']);
    }

    // Authorization uses the real logged-in user, never an ID the model supplied.
    $order = $user->orders()->findOrFail($toolCall['args']['order_id']);

    return $this->tools->execute($toolCall['name'], $order);
}

7. Validate AI output like any untrusted input

If your code consumes AI output (JSON, SQL, a category, a price), parse and validate it exactly as you would a form submission. The model will sometimes return extra prose, a wrong type, a missing field, or a confident wrong answer.

$raw = $llm->ask($prompt . "\nReturn ONLY JSON: {\"category\": string, \"priority\": 1-5}");

$data = json_decode($raw, true);

$validator = Validator::make($data ?? [], [
    'category' => ['required', Rule::in(['billing', 'bug', 'feature', 'other'])],
    'priority' => ['required', 'integer', 'between:1,5'],
]);

if ($validator->fails()) {
    // Retry a limited number of times, then fall back: never loop forever.
    return $this->fallbackToHumanTriage($ticket);
}

$ticket->update($validator->validated());

Notice three habits: an allow-list for categories, a range check for numbers, and a bounded fallback instead of trusting (or endlessly retrying) the model. Most providers also offer a structured output / JSON-schema mode. Use it, and still validate.

Test it too. AI features need tests just like anything else. Keep a small “golden set” of inputs with expected outputs and run it in CI whenever you change the prompt or model, so you catch regressions instead of users.


8. Good prompts are specs, and specs reduce cost

A vague prompt gets a vague answer, which leads to a follow-up, which leads to another. Every retry re-sends context and burns tokens. A precise prompt is the cheapest optimization you have.

Vague (costly) Specific (cheap)
“Fix my code” “This Laravel 11 action throws N+1 on orders->items. Fix it with eager loading. Return only a diff.”
“Write tests” “Write Pest tests for RefundService::issue() covering: full refund, partial refund, and already-refunded. No DB, mock the gateway.”
“Make it better” “Reduce cyclomatic complexity. Keep the public signature and behaviour identical.”

A reliable prompt has five parts: role/context, task, constraints, examples, and output format.

Context:     Laravel 11, PHP 8.3, Pest for tests.
Task:        Refactor OrderController@store into a service class.
Constraints: Keep route and response shape unchanged. No new packages.
Example:     Follow the style of app/Services/InvoiceService.php.
Output:      A unified diff only, no explanations.

The “Output” line matters most for cost: asking for only the diff can cut output tokens dramatically versus a full rewrite with commentary.


9. Control cost before it controls you

AI cost is usage-based, so a bug or a loop can become a bill. Four levers handle most of it:

  1. Right-size the model. Use a small, cheap model for classification and formatting. Save the large one for hard reasoning.
  2. Cap everything. Set max_tokens, request timeouts, retry limits, and a per-user or per-day budget.
  3. Cache. Identical questions shouldn’t be paid for twice. Many providers also discount repeated prompt prefixes (“prompt caching”), so keep the stable part of your prompt first.
  4. Send less. Trim history, retrieve only relevant chunks, and ask for concise output.
public function classify(string $text, User $user): string
{
    // 1. Budget guard: a runaway loop hits a wall, not your credit card.
    if ($this->usage->todayFor($user) > config('ai.daily_token_limit')) {
        throw new BudgetExceededException();
    }

    // 2. Cache: same input, same answer, zero tokens.
    $key = 'ai:classify:' . sha1($text);

    return Cache::remember($key, now()->addDay(), function () use ($text, $user) {
        // 3. Route to the cheap model and cap the output.
        $result = $this->llm->ask(
            model: config('ai.models.cheap'),
            prompt: "Classify as billing|bug|feature|other. Reply with one word.\n\n{$text}",
            maxTokens: 5,
        );

        $this->usage->record($user, $result->tokensIn, $result->tokensOut);

        return trim($result->text);
    });
}

Real-world failure: the silent retry loop

A very common pattern: an agent or script calls a model, gets a malformed answer, retries with the whole conversation appended, fails again, and repeats overnight. Nothing crashes. Nothing alerts. The bill just grows. The fix is boring and effective: a retry cap, a per-job token budget, and an alert when daily spend passes a threshold.

Measure first. Log tokens in, tokens out, model, latency, and feature name for every call. You can’t optimize what you can’t see.


10. You own the code. AI is a tool, not a teammate who takes the blame

If you commit it, it’s yours. “The AI wrote it” isn’t a defense in a code review, an outage, or a security audit. AI generates code faster than people can review it, which means technical debt can now accumulate at machine speed: duplicated logic, inconsistent patterns, untested branches, and code nobody on the team truly understands.

Keep your standards exactly where they were:

  • Understand before you merge. If you can’t explain a line, you can’t maintain it. Ask the AI to explain it, then verify.
  • Keep changes small. Small AI-assisted PRs are reviewable. A 3,000-line generated diff is not.
  • Write the tests first (or demand them). Tests turn “looks right” into “is right”.
  • Follow your architecture. Tell the AI your patterns. Don’t let it invent a new one per file. See Must-Know Code Quality Practices and AI-Assisted Development in Fintech Engineering.
  • Keep learning the fundamentals. Juniors who skip understanding now become seniors who can’t debug later.

The 10-point checklist

Before you ship anything AI-assisted, ask yourself:

  1. Did I verify facts and APIs instead of trusting confident output?
  2. Do I know roughly how many tokens this call costs?
  3. Am I sending only the context the model needs?
  4. Did I review the generated code like a stranger’s PR (security, edge cases)?
  5. Are all suggested dependencies real, maintained, and audited?
  6. Are secrets and personal data redacted or excluded?
  7. Can untrusted input reach the model, and can the model reach anything dangerous?
  8. Is AI output validated against a schema or allow-list?
  9. Are there max_tokens, retry caps, caching, and a spend limit?
  10. Do tests cover it, and can I explain every line I’m committing?

Key takeaways

  • LLMs predict plausible text; they don’t guarantee truth. Ground and verify.
  • Tokens and context are your budget. Measure them, trim them, cap them.
  • AI code and AI output are untrusted input. Review, validate, and apply least privilege.
  • Real failures (Air Canada, Samsung, hallucinated packages, the deleted production database) came from ordinary gaps: no grounding, leaked data, no review, too much access.
  • You stay accountable. AI raises your speed. Your judgment still sets the quality.

AI won’t replace engineers who understand it, but it will amplify the habits they already have, good or bad. Pick one item from the checklist and apply it in your next pull request.