Cloud Native AI Summit
All speakers
Renaldi Gondosubroto

Speaker

Renaldi Gondosubroto

Director · Cloudetica Solutions

About

Renaldi Gondosubroto is the Director of Cloudetica Solutions, where he leads the building of open-source solutions on cloud platforms. He brings over a decade of software development experience and has been active within the research community, putting a lot of his research focus within IoT and virtual reality. He is the author of three technical books, including "Using Amazon Bedrock: Learn to Architect, Secure and Optimize Generative AI Applications on AWS" published by Wiley. Having spoken at over 50 events and conferences, he has been an international speaker for the past six years, sharing his experiences and projects. He also currently is an AWS Subject Matter Expert (SME) for its Professional, Associate and Specialty Certifications and holds all 14 AWS certifications. He aims to build open-source solutions which can both help people achieve more value in what they do and promote best practices for fellow developers.

Session

The Build Passed and the Model Failed

Talk

A platform team had reached the point where changes to its AI application looked reassuringly similar to changes anywhere else in the software estate. Prompts lived in Git, model configuration was versioned, pull requests required review, automated tests ran on every change, and deployments moved through the normal delivery pipeline. One release passed every check and deployed without an infrastructure error. The application was still worse than the version it replaced. Nothing had crashed. Latency remained within its target and the APIs were healthy. The regression was behavioural. A relatively small change to the system prompt had improved performance on one class of requests while making another less reliable. The existing pipeline had been built to detect software failures, so from its perspective the release was completely healthy. This session follows the engineering response to that incident and the architecture that emerged from it. Instead of placing AI evaluation outside the delivery process as a periodic benchmarking exercise, the team made model behaviour part of the release contract. Prompt changes, model upgrades, retrieval configuration and tool definitions became versioned application dependencies with their own validation requirements. We will walk through the resulting delivery architecture, including representative evaluation datasets, deterministic checks around tool use and structured outputs, model-scored evaluation where deterministic assertions are insufficient, and regression thresholds that can prevent a change from progressing. The case study also covers how those evaluations were kept useful as the application changed, rather than allowing the test set to become a static collection that the system gradually learned to satisfy. The production side required a different set of controls. Offline evaluation reduced the number of bad releases, but it could not reproduce every real request. The deployment process therefore introduced shadow evaluation, limited canary releases and production feedback signals so that behavioural regressions could be detected without making every model change an all-or-nothing deployment. A significant part of the session focuses on the engineering problems that appear once this becomes CI/CD. Model outputs are not always deterministic, evaluation adds cost and latency to pipelines, and an aggregate score can hide a serious regression affecting a small but important class of requests. The architecture has to account for those properties instead of pretending an LLM test suite behaves like ordinary unit tests. The case study ultimately changed the team's definition of a successful deployment. Healthy infrastructure was no longer enough. An AI release also had to demonstrate that the behaviour users depended on had not materially regressed. Attendees will leave with a practical architecture for bringing prompts, models and behavioural evaluation into a cloud-native delivery pipeline without turning every pull request into a manual AI review.

Speaking at

  • Cloud Native AI Summit — Melbourne

    October 28–29, 2026

    View event →