Welcome to the Sapiver Forge Weekly AI Briefing, where we turn noisy AI news into clear, usable learning. Human-led. AI-empowered. This week’s central story is not one model beating another on a chart. It is AI moving deeper into the workflow itself. The clearest example came from Canva. Canva AI 2.0 was announced as a research preview, and the ambition is much larger than generating a picture. Canva says the new system can work conversationally, keep designs layered and editable, connect to tools such as Gmail, Slack, Google Drive and Notion, research the web, schedule recurring work, apply brand rules, build spreadsheets and create interactive experiences. That sounds powerful. It also changes the risk. A tool that makes one image can make one bad image. A connected system with access to email, files, calendars and publishing can carry one bad assumption much further. So the Sapiver Forge verdict is test carefully. Start with one draft output. Give it approved source material. Do not let it publish automatically. Check what every connector can access, and remove permissions you do not need. For a creator, this could mean turning an approved campaign brief into a draft social graphic, a presentation and a simple landing page without rebuilding the same idea three times. For a small business, it could mean summarising customer feedback, organising the themes in a sheet and preparing a draft response plan. In both cases, the value comes from continuity between steps. The human job is to define the source of truth, protect confidential material and check that the final claim, price, date and call to action are correct. The second major development was Inkling from Thinking Machines Lab. Inkling is an open-weight multimodal model. In plain English, the learned parameters are available for others to customise or deploy, subject to the licence and practical hosting requirements. The company describes a very large mixture-of-experts model with a million-token context window and support for text, images and audio. This matters because open weights can reduce dependence on one hosted provider and allow deeper customisation. But open does not mean effortless. Hosting costs, security, monitoring, evaluation and incident response move toward the deployer. For most creators and small businesses, Inkling is one to watch. The useful question is not whether you can fine-tune a model. It is whether you have a repeatable problem that simpler prompting and document retrieval cannot solve. Here is a practical way to separate curiosity from a genuine use case. Write down the job you want the model to perform, the data it would need, the acceptable response time and the consequence of a wrong answer. If the job is occasional, low-risk and already served by a hosted model, self-deployment may add work without adding value. If the job is frequent, specialised or constrained by privacy and control requirements, an open-weight option becomes more interesting. That is a business decision before it is a technical one. GitHub also moved AI into a decision point. Its public preview of AI-powered security detections places labelled findings directly in pull requests. The detections run when a pull request is opened or updated, and GitHub says they are informational rather than merge-blocking. Access requires GitHub Code Security, CodeQL setup, enterprise permission, a Copilot licence and AI credits. The wider lesson applies beyond software. Put AI review where a human decision already happens. For code, that is the pull request. For a newsletter, it is the final send review. For a product page, it is the claims and price check before publication. AI can flag a concern, but a person should still examine the evidence. The label is important as well. An AI-generated security finding is a lead, not a verdict. It can be wrong, incomplete or focused on the visible symptom rather than the root cause. A sensible review records whether the issue was confirmed, what evidence supported the decision and whether a broader test is needed. That creates feedback for the team instead of quietly accepting or dismissing the machine's suggestion. Google’s NotebookLM was renamed Gemini Notebook and gained a more operational capability: code execution inside a secure cloud computer. Google says this is available now for certain Ultra and Workspace customers and will roll out to Pro users on the web over the coming weeks. That staged rollout matters. An announcement is not the same as access on every account. For eligible users, the useful experiment is small. Upload an anonymous, approved data set and its category definitions. Ask the notebook to calculate counts and percentages, show the method, flag unmatched rows and create a chart. Then manually check a sample. Code that runs can still answer the wrong question, so source grounding and calculation review both matter. Keep the first test deliberately boring. A monthly content log or a table of anonymised enquiries is better than a sensitive customer database. Compare the notebook's totals with a calculation you already trust. If the numbers disagree, investigate the definition, filtering and missing-data rules before adding more automation. A successful demonstration is not proof that every future input will behave the same way. OpenAI published two developments that are less like products and more like operating lessons. The first is a scorecard for judging AI by completed work. The practical measure is cost per successful task. Add the model or subscription cost, staff time, human review, retries and rework. Then divide that total by the number of outputs that actually met the quality bar. This is a vendor framework, not a neutral standard, but the method is useful with any provider. Try it on one task that happens every week. Define what done means before the test. Run ten normal examples, not ten easy ones. Count accepted results, corrections and failures. A cheap model that needs three retries may cost more than a stronger model that succeeds once. Equally, an impressive model may still be poor value if every result needs heavy editing. The second OpenAI development is GPT-Red, an internal automated red-teaming system. OpenAI says it used GPT-Red to find vulnerabilities and adversarially train GPT-5.6 against prompt injection. Prompt injection is an instruction hidden in something the AI reads, such as a webpage, email, file or tool response. It tries to redirect the system, reveal information or make it take an unsafe action. Automated red-teaming can generate many attack variations faster than people can create them manually. GPT-Red is not a customer tool you can simply switch on. The practical lesson is to treat external content as untrusted. Separate research from action. If an AI can read the web, it should not automatically gain permission to send a payment, publish a post, delete a file or upload private data. Keep a named human at those boundaries. Consumer safety also moved forward. OpenAI said it is expanding teen safeguards and rolling out age prediction on consumer plans. The aim is to estimate whether an account may belong to someone under eighteen and apply a safer experience. That may reduce exposure to sensitive content, but it also creates questions about privacy, mistakes and appeals. If your own product may reach younger users, do not copy the technology blindly. Start with clear age expectations, minimal data, visible safety settings and a route to a human. Regulation is shaping the distribution layer too. The European Commission issued binding measures intended to give competing AI assistants access to Android features used by Google’s own services, while also setting conditions for eligible search providers and AI chatbots to receive anonymised Google search data. The immediate work falls on Google and qualifying providers. For normal users, the important point is future choice rather than a setting to change today. Reuters reported implementation milestones beginning in January twenty twenty-seven, with some Android user benefits from July twenty twenty-seven. Those dates should be rechecked nearer the time because implementation can change. Finally, the infrastructure buildout remains active. ASML raised its annual outlook and described plans to expand advanced chipmaking-system capacity. That does not mean every AI company will succeed, and it is not investment advice. It does mean small teams should treat AI as a recurring operating cost. Link every subscription to a workflow, set spending limits, and keep a fallback for anything important. Here is the one experiment to take away this week. Choose one AI-assisted task. Write down what goes in, what the AI may do, what it must never do, what a successful result looks like, and who approves it. Run five to ten real examples. Count the cost, retries and correction time. Then make one decision: keep it, tighten it, pause it or remove it. Use a simple review card for each attempt. Record whether the first answer was usable, how many minutes a person spent checking it, what had to be corrected and whether any private or unapproved information appeared. At the end, do not average away a serious failure. One fabricated price, exposed record or unauthorised action can matter more than several fast successes. The goal is dependable assistance, not a flattering demonstration. Next week, watch for clearer Canva access details, wider Gemini Notebook code execution, early independent testing of Inkling, and practical implementation information around the European Commission’s measures. Sapiver Forge does not ask you to chase every release. We help you turn change into one useful, human-led system. Visit the Sapiver Forge website for the full learning guide, daily briefings and practical exercises. You can also follow Sapiver Forge AI Briefing in your usual podcast app, including Spotify. Human-led. AI-empowered. Turning human input into clear, usable systems.