https://raw.githubusercontent.com/ajmaradiaga/feeds/main/scmt/topics/SAP-HANA-Cloud-blog-posts.xml SAP Community - SAP HANA Cloud 2026-07-17T20:01:05.573359+00:00 python-feedgen SAP HANA Cloud blog posts in SAP Community https://community.sap.com/t5/technology-blog-posts-by-sap/intel-and-sap-deepen-collaboration-for-sap-hana-cloud-co-development/ba-p/14419673 Intel and SAP Deepen Collaboration for SAP HANA Cloud Co-Development Initiative 2026-06-16T16:30:56.507000+02:00 stefan_baeuerle https://community.sap.com/t5/user/viewprofilepage/user-id/512708 <H1 id="toc-hId-1688357658">A new chapter of co-engineering around SAP HANA Cloud, focused on cloud-native scale, performance, and what’s next for customers.</H1><P class="lia-align-right" style="text-align : right;"><BR /><U>Disclaimer:</U> This blogpost has been co-published at&nbsp;<A href="https://community.intel.com/t5/Blogs/Tech-Innovation/Cloud/Intel-SAP-put-HANA-Cloud-Platform-at-the-center-of-a-newly/post/1751411" target="_blank" rel="noopener nofollow noreferrer">community.intel.com.</A></P><H2 id="toc-hId-1620926872">Introduction</H2><P>Intel and SAP are expanding their long-standing collaboration through a dedicated SAP HANA Cloud co-development initiative, bringing engineering teams closer together to help SAP HANA Cloud run even better on current and future Intel platforms. This is the start of a new chapter in how Intel and SAP co-innovate: aligning SAP’s product direction with Intel’s platform roadmap, so optimizations show up where they matter most, in real customer environments running on leading cloud service providers.<BR /><BR /></P><H2 id="toc-hId-1424413367">A new chapter of co-engineering for SAP HANA Cloud</H2><P>This initiative is founded on the principle that early and comprehensive collaboration between platform and software engineering teams leads to enhanced customer experiences. The ongoing partnership has demonstrated great benefits for SAP HANA during the last decade.</P><P>This collaboration creates a direct technical bridge between SAP product development and Intel architecture/platform roadmaps, helping ensure new capabilities and platform features and enhancements translate into real, production-ready outcomes across SAP's certified cloud service providers.<BR /><BR /></P><H2 id="toc-hId-1227899862">Value Across the Ecosystem</H2><P>The co-development initiative is designed to deliver clear value across all stakeholders:</P><P><STRONG>For SAP:</STRONG> Architectural optimizations for SAP HANA Cloud features, accelerated technology adoption, and reduced risk during transitions, such as migrating to new Intel Xeon generations, introduce new and/or enhanced Intel features.</P><P><STRONG>For Intel:</STRONG> Ensuring SAP HANA Cloud runs at its best on Intel x86 Architectures and maintains and improves competitiveness against alternative architectures. Strengthening Intel's position for SAP's strategic cloud deployments at the Cloud Service Providers (CSPs).</P><P><STRONG>For Customers and CSPs:</STRONG> Improved price-performance, efficiency, scalability, and resiliency are direct benefits from joint SAP-Intel developments in both performance and feature sets.<BR /><BR /></P><H2 id="toc-hId-1031386357">What we’re building together</H2><P>Intel and SAP share a vision to proactively optimize SAP HANA Cloud on current and future Intel platforms, delivering value for SAP, cloud service providers, and, most importantly, customers running mission-critical workloads on SAP HANA Cloud. By connecting SAP product development with Intel architecture and platform roadmaps, the teams are creating an engine for continuous optimization, early validation, and faster production readiness.<BR /><BR /></P><H2 id="toc-hId-834872852">Where we’ll focus first</H2><P>The initiative is organized around three practical workstreams designed to turn joint engineering into customer-impacting outcomes:</P><H4 id="toc-hId-896524785">Platform Development</H4><P>Intel and SAP will work together on feature and workload optimization across key SAP HANA Cloud scenarios, such as OLTP, OLAP, scale-up, and scale-out, while also streamlining the introduction of new Intel® Xeon® processor-based cloud instances and their new capabilities.</P><H4 id="toc-hId-700011280">Performance Benchmarking</H4><P>The teams will jointly review SAP HANA Cloud–specific benchmarks and tune for strong, repeatable results on Intel platforms at CSPs, addressing constraints observed in production and helping ensure competitive performance compared with alternative architectures. This applies to Transactional, Analytical and AI-enhanced workloads.</P><H4 id="toc-hId-503497775">Product Readiness</H4><P>Intel will help SAP accelerate the adoption of new instance types and configurations across SAP HANA Cloud editions and ensure that proven optimizations carry forward smoothly over time, so customers benefit consistently throughout the service lifecycle.</P><P>&nbsp;</P><H2 id="toc-hId-48818832">What does this mean for customers and cloud service providers for SAP solutions?</H2><P>For customers, this initiative is about turning platform innovation into practical outcomes, better price-performance, better efficiency, and faster readiness for what comes next. For cloud service providers, it’s about tighter alignment and earlier validation so SAP HANA Cloud can take advantage of new Intel capabilities in real production environments.<BR /><BR /></P><H2 id="toc-hId-199559684">Looking ahead: stay tuned</H2><P>As this co-development effort ramps up, SAP HANA Cloud is positioned to benefit from optimizations on Intel Xeon platforms, including Intel Xeon 6 processors and future generations. The focus is straightforward: accelerate production readiness, scale and efficiency, and deliver continuous improvements that customers can use, without disruption.</P><P>This is just the beginning. Intel and SAP are deepening co-engineering for SAP HANA Cloud and Cloud Native Architectures at SAP to help SAP to scale, perform, and evolve on Intel platforms as SAP Cloud ERP adoption accelerates. Stay tuned for the innovations Intel and SAP will deliver to customers, and for more updates as the teams turn this initiative into tangible improvements across cloud environments.</P><P>&nbsp;</P><P>&nbsp;</P><P>&nbsp;</P><P><STRONG>Notices and Disclaimers</STRONG></P><P>Performance varies by use, configuration, and other factors. Learn more on the <A href="https://edc.intel.com/content/www/us/en/products/performance/benchmarks/overview/" target="_blank" rel="noopener nofollow noreferrer">Performance Index site</A>.&nbsp;</P><P>Performance results are based on testing as of dates shown in configurations and may not reflect all publicly available ​updates.&nbsp; See backup for configuration details.&nbsp; No product or component can be absolutely secure.</P><P>Your costs and results may vary.</P><P>Intel technologies may require enabled hardware, software, or service activation.</P><P>© Intel Corporation.&nbsp; Intel, the Intel logo, and other Intel marks are trademarks of Intel Corporation or its subsidiaries.&nbsp; Other names and brands may be claimed as the property of others.</P> 2026-06-16T16:30:56.507000+02:00 https://community.sap.com/t5/data-professionals-blog-posts/working-with-bdc-in-the-sap-business-ai-platform-the-data-analyst/ba-p/14422275 Working with BDC in the SAP Business AI Platform: The Data Analyst Perspective 2026-06-18T10:44:12.492000+02:00 ThierryAudas https://community.sap.com/t5/user/viewprofilepage/user-id/8851 <P>As a data analyst working with SAP Business Data Cloud (BDC), you already rely on governed data models, consistent KPIs, and trusted semantics to produce analytics and insights. The Autonomous Enterprise strategy with the SAP Business AI Platform (BAIP) doesn’t change that foundation, it extends it.</P><P>What <EM>does</EM> change is how insights are generated, explained, and acted on. Analytics is no longer the final step in the data value chain. It becomes an active input into AI‑driven reasoning and execution, grounded in the same business context that you define and validate today.&nbsp;At the core of this shift is how BDC participates in the <EM>Contextualize and Reason</EM>&nbsp;services of BAIP, alongside the SAP Knowledge Graph. BDC provides the harmonized data foundation and business semantics, while the SAP Knowledge Graph encodes relationships and business context. Together, they form the trusted context layer that AI agents rely on to reason and act.</P><H3 id="toc-hId-1947234238"><STRONG>What stays the same</STRONG></H3><P>Your core analytical responsibilities remain intact.&nbsp;You continue to work with semantic models defined in SAP Datasphere, SAP BW, and SAP HANA Cloud, accessing SAP and non-SAP data, where business meaning is explicitly modeled. These models, together with master data governed through SAP Master Data Governance, form the structured input to the knowledge graph that AI relies on. SAP Analytics Cloud (SAC) remains your primary environment for analysis, visualization, and planning.</P><P>The difference is that these analytical assets are no longer only consumed by humans, they are part of the contextual foundation used by AI agents.</P><H3 id="toc-hId-1750720733"><STRONG>What changes in practice</STRONG></H3><P>The most important shift is that AI now consumes the same semantic and contextual layer you work with.&nbsp;In the Autonomous Enterprise, AI agents do not generate insights from isolated datasets. They reason over BDC data products enriched with relationships captured in the SAP Knowledge Graph. When an AI agent produces a narrative, explains a deviation, or recommends an action, it operates on top of business objects, hierarchies, and dependencies that you already know.</P><P>This changes the validation process. Instead of questioning the origin of an AI insight, you can trace it back to governed data products and their relationships within the broader business context.&nbsp;Analytics shifts from being a downstream consumer of data to being a shared semantic contract between humans and AI.</P><H3 id="toc-hId-1554207228"><STRONG>How AI</STRONG><STRONG>‑</STRONG><STRONG>assisted analytics works with BDC</STRONG></H3><P>AI‑assisted insight generation doesn’t bypass your analytical models. It builds on them.&nbsp;Data products defined in BDC, whether sourced from SAP applications, BW objects, or Datasphere models, carry business meaning, lineage, and access rules. BAIP uses these data products as the context layer for AI reasoning. When AI generates narratives, summaries, or explanations, it does so against this governed semantic layer.</P><P>For you as a data analyst, this means faster exploration and explanation without sacrificing consistency. You spend less time reconciling definitions and more time interpreting results.</P><H3 id="toc-hId-1357693723"><STRONG>Impact on trust and validation</STRONG></H3><P>One of the biggest challenges with AI‑generated insights is trust. With BAIP, trust is achieved by design.&nbsp;Because AI agents rely on the same BDC‑managed semantics you use, AI outputs inherit the same KPI logic, hierarchies, and filters. Access controls and data policies are enforced consistently, whether the consumer is a dashboard, a planner, or an AI agent.</P><P>This dramatically reduces the need for parallel validation loops and manual reconciliation.</P><H3 id="toc-hId-1161180218"><STRONG>How this maps to BDC components</STRONG></H3><P>For data analysts, the component mapping is straightforward.</P><UL><LI>SAP Datasphere continues to be the primary place where analytical semantics are modeled and exposed as data products.</LI><LI>SAP BW continues to provide enterprise‑grade, historically rich business logic that remains authoritative for many KPIs.</LI><LI>SAP HANA Cloud supports advanced analytical queries and high‑performance access to data products.</LI><LI>SAP Analytics Cloud remains the main consumption layer for analysis, planning, and increasingly AI‑assisted insight generation.</LI><LI>SAP Master Data Governance ensures that master data consistency underpins all analytical and AI‑driven outcomes.</LI></UL><P>Together, these components form the trusted analytical foundation that AI agents consume through BAIP.</P><H3 id="toc-hId-964666713"><STRONG>What to focus on now</STRONG></H3><P>As a data analyst, the most valuable preparation is not learning new AI tools. It is strengthening the analytical foundation AI will rely on.&nbsp;Focus on well‑defined KPIs, stable semantics, clear lineage, and reusable data products. The better your analytical models are today, the more reliable AI‑assisted insights will be tomorrow.</P><H3 id="toc-hId-768153208"><STRONG>The takeaway</STRONG></H3><P>BAIP doesn’t replace analytics with AI. It connects analytics and AI through a shared semantic foundation.&nbsp;For data analysts, this means higher leverage: the models you build and validate do not just inform people, they inform AI‑driven decisions as well.</P><P>Join the Data Professionals community discussions to share analytical patterns, discuss AI‑assisted insights, and explore how others are evolving their BDC analytics foundations in the Autonomous Enterprise era.</P><P><EM>This post is a follow up to the previous post <A href="https://community.sap.com/t5/data-professionals-blog-posts/from-ai-in-applications-to-ai-on-applications/ba-p/14415270" target="_blank">From ‘AI in Applications’ to ‘AI on Applications’</A>. It is the first in a series that explores how the SAP Business AI Platform (BAIP) impacts day‑to‑day work across key BDC personas, starting with the data analyst perspective.</EM></P> 2026-06-18T10:44:12.492000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/ba-p/14411326 Hands-on Tutorial: Vibe coding SAP HANA Cloud Machine Learning with BAS, Python, Claude and Cline 2026-06-18T11:12:53.472000+02:00 AndreasForster https://community.sap.com/t5/user/viewprofilepage/user-id/14188 <P>Have a Large Language Model write the Machine Learning code for you next project!</P><P>Sit back and watch the <STRONG>SAP Business Application Studio (BAS)</STRONG> use <STRONG>Anthropic&nbsp;Claude</STRONG> to create <STRONG>Jupyter Notebooks</STRONG> and write the Python code to trigger the <STRONG>Predictive Analysis Library</STRONG> in <STRONG>SAP HANA Cloud (or SAP Datasphere)</STRONG>. This example uses Claude, but there is a long list of models to choose from through <STRONG>SAP AI Core</STRONG>. In this tutorial we go through a time-series forecast, but that's just an example, you can use this ability to vibe code in BAS for all sorts of requirements.</P><P>It's fantastic to see how the code is written automatically, and how good the code is! But to keep it real, work with that automatically created code as if it had been written by a human intern. Inspect the code to ensure it does what you want it to do.</P><P><ul =""><li style="list-style-type:disc; margin-left:0px; margin-bottom:1px;"><a href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/ba-p/14411326#toc-hId-1688116295">Architecture</a></li><li style="list-style-type:disc; margin-left:0px; margin-bottom:1px;"><a href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/ba-p/14411326#toc-hId-1491602790">Prerequisites</a></li><li style="list-style-type:disc; margin-left:0px; margin-bottom:1px;"><a href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/ba-p/14411326#toc-hId-1295089285">Setup</a></li><li style="list-style-type:disc; margin-left:0px; margin-bottom:1px;"><a href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/ba-p/14411326#toc-hId-1098575780">Vibe coding test</a></li><li style="list-style-type:disc; margin-left:0px; margin-bottom:1px;"><a href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/ba-p/14411326#toc-hId-902062275">Adding the latest documentation</a></li><li style="list-style-type:disc; margin-left:0px; margin-bottom:1px;"><a href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/ba-p/14411326#toc-hId-705548770">Vibe code HANA Machine Learning</a></li><li style="list-style-type:disc; margin-left:0px; margin-bottom:1px;"><a href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/ba-p/14411326#toc-hId-509035265">Summary</a></li></ul></P><P>Big thanks go to&nbsp;<a href="https://community.sap.com/t5/user/viewprofilepage/user-id/862667">@sherwin_emami</a>&nbsp;for introducing me to Cline and for showing me the ropes of vibe coding in BAS!</P><P>&nbsp;</P><H1 id="toc-hId-1688116295">Architecture</H1><P>The architecture centers around the SAP Business Application Studio, where you can already do you Python scripting in a Jupyter Notebook by hand.&nbsp;</P><P>The coding agent <A href="https://cline.bot/" target="_blank" rel="noopener nofollow noreferrer">Cline</A> is installed as extension in BAS. Cline connects to <A href="https://help.sap.com/docs/sap-ai-core/sap-ai-core-service-guide/what-is-sap-ai-core" target="_blank" rel="noopener noreferrer">SAP AI Core</A> to integrate the Large Language Models (LLM) that process the user requests and produce the code. You can choose which LLM you want to use. In this example we use Claude, which is hosted outside the SAP environment. Other models, i.e. from Mistral, are hosted on SAP infrastructure.</P><P>SAP HANA Cloud holds the data and Machine Learning algorithms. The same setup would also work with the SAP HANA Cloud that is embedded in SAP Datasphere.</P><P>Optionally, you can also add th<FONT color="#000000">e <A href="https://github.com/upstash/context7" target="_blank" rel="noopener nofollow noreferrer">Context7</A> MCP server&nbsp;to&nbsp;</FONT>Cline to get access to additional / current software documentation.</P><P>You may have heard of the SAP Business AI Platform, which is also shown in the architecture diagram. However, here we use the same components that you know from the SAP Business Technology Platform (BTP). You can implement this scenario on any BTP environment.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="010 Architecture.png" style="width: 687px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418122iF4B2BF5DC040C85B/image-size/large?v=v2&amp;px=999" role="button" title="010 Architecture.png" alt="010 Architecture.png" /></span></P><P>&nbsp;</P><H1 id="toc-hId-1491602790">Prerequisites</H1><P>You will need the following components to implement this blog's scenario:</P><UL><LI>SAP Business Application Studio, configured to run Jupyter Notebooks, see this <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/setting-up-python-in-sap-business-application-studio-to-trigger-hana-cloud/ba-p/14387837" target="_blank">blog</A></LI><LI>SAP HANA Cloud<UL><LI>configured to run the Predictive Analysis Library, see the same&nbsp;<A href="https://community.sap.com/t5/technology-blog-posts-by-sap/setting-up-python-in-sap-business-application-studio-to-trigger-hana-cloud/ba-p/14387837" target="_blank">blog</A></LI><LI>with the table HOTELNIGHTS from <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-machine-learning-with-sap-hana-cloud-and-bas/ba-p/14404096" target="_blank">this other blog&nbsp;</A>uploaded</LI></UL></LI><LI>SAP AI Core, on an <A href="https://help.sap.com/docs/sap-ai-core/sap-ai-core-service-guide/service-plans?version=CLOUD" target="_blank" rel="noopener noreferrer">Extended Service Plan</A>. SAP AI Launchpad is not needed for the vibe coding.</LI></UL><P>&nbsp;</P><H1 id="toc-hId-1295089285">Setup</H1><P>Install the Cline extension into the BAS. This should take only a few seconds.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="020 Cline install.gif" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418212i066849D072ADF770/image-size/large?v=v2&amp;px=999" role="button" title="020 Cline install.gif" alt="020 Cline install.gif" /></span></P><P>&nbsp;</P><P>Then get a Service Key from your SAP AI Core instance.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="030 BTP cockpit service key.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418213i236AFDBC849D254F/image-size/large?v=v2&amp;px=999" role="button" title="030 BTP cockpit service key.png" alt="030 BTP cockpit service key.png" /></span></P><P>&nbsp;</P><P>Continue in the Cline extension and give it access to SAP AI Core.</P><UL><LI>Select the option "Bring my own API key".</LI><LI>From the "API Provider" dropdown select "SAP AI Core" and enter the values from your SAP AI Core Service Key:<UL><LI><STRONG>AI Core Client ID:</STRONG> value of <STRONG>clientid</STRONG></LI><LI><STRONG>AI Core Client Secret:</STRONG> value of&nbsp;<STRONG>clientsecret</STRONG></LI><LI><STRONG>AI Core Base URL:</STRONG> value of&nbsp;<STRONG>AI_API_URL</STRONG></LI><LI><STRONG>AI Core Auth URL:</STRONG> value of <STRONG>url</STRONG></LI><LI><STRONG>AI Core Resource Group:</STRONG> By default the Resource Group is "default", unless you created your own in SAP AI Core</LI></UL></LI><LI>Tick "Orchestration Mode".&nbsp; This makes all applicable models from SAP AI Core available, without having to deploy them first</LI><LI>Select your preferred model. Here I am going with Claude 4.7 Opus.</LI></UL><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="040 Cline config.png" style="width: 523px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418302i247BDB8731227726/image-size/large?v=v2&amp;px=999" role="button" title="040 Cline config.png" alt="040 Cline config.png" /></span></P><P>&nbsp;</P><P>Click "Continue" and Cline is asking what it can do your you&nbsp;</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="050 Cline.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418325i50BAC6C96E885AE8/image-size/large?v=v2&amp;px=999" role="button" title="050 Cline.png" alt="050 Cline.png" /></span></P><P>&nbsp;</P><H1 id="toc-hId-1098575780">Vibe coding test</H1><P>Use a simple task to get familiar with Cline. It's important to know the distinction between "Plan" and "Act" mode.&nbsp;</P><P><STRONG>Plan mode:</STRONG> Typically it is best to start in Plan mode. You describe what you would Cline built for you and Cline is building a plan for achieving this. This can be a discussion, which shapes how to address this task. In this phase, Cline can access files (if you grant access), but it cannot make any changes to your existing file or create new files.</P><P><STRONG>Act mode: </STRONG>Once you are happy with the plan you can switch to the "Act" mode, in which Cline executes the plan. Here Cline can, for instance, create new files or modify existing ones.</P><P>Let's try it out!</P><P>First select the folder in which you want to work, ie "<SPAN>home/user/projects/hanaml". If you are unsure about the folder, see <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/setting-up-python-in-sap-business-application-studio-to-trigger-hana-cloud/ba-p/14387837" target="_blank">this blog</A>.</SPAN></P><P>Make sure that Cline is in "Plan" mode (the little button below the prompt window). Check which permissions you want to give Cline. Then enter this task and Cline is creating the plan to achieve this.</P><TABLE border="1" width="100%"><TBODY><TR><TD width="100%"><FONT color="#3366FF"><EM>Create a jupyter notebook that contains a greeting</EM></FONT></TD></TR></TBODY></TABLE><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="060 vibe test plan.gif" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418339iB025DB9B0E3071E5/image-size/large?v=v2&amp;px=999" role="button" title="060 vibe test plan.gif" alt="060 vibe test plan.gif" /></span></P><P>&nbsp;</P><P>We are happy with the plan and want this executed. Click on the "Act" tab. You will see how Cline is going through the plan. If you are using the default permissions, you will be asked to confirm the creation of the Notebook file and indeed the file opens up!</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="070 vibe test act.gif" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418340i5DFC1E1599EA1E1D/image-size/large?v=v2&amp;px=999" role="button" title="070 vibe test act.gif" alt="070 vibe test act.gif" /></span></P><P>&nbsp;</P><P>You can run the code in the notebook. And the file is there as requested.&nbsp;</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="080 vibe test run.gif" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418341i74A040B119806D1A/image-size/large?v=v2&amp;px=999" role="button" title="080 vibe test run.gif" alt="080 vibe test run.gif" /></span></P><P>&nbsp;</P><P>Mission accomplished, the vibe coding test was successful!</P><P>&nbsp;</P><P>A few hints in case this didn't work for you:</P><UL><LI>With Cline's default settings you will probably get a few approval requests. You need to confirm these for the process to continue.</LI><LI>These requests are not always immediately shown on screen. Check whether you need to scroll further down in the Cline window to see the latest output.</LI><LI>Sometimes Cline seems to hang. Refreshing the browser seems to fix it for me</LI><LI>Sometimes Cline opening a file results in the error "The editor could not be opened due to an unexpected error. Please consult the log for more details.". Clicking the&nbsp;"Try again" usually fixes it for me.</LI><LI>You know that the "Act" mode has completed, when you see "Start New Task".</LI></UL><P>&nbsp;</P><H1 id="toc-hId-902062275">Adding the latest documentation</H1><P>Before creating code for your business, let's make Cline even more useful. Currently the Large Language Model that we are using might not know the latest documentation of the libraries that we want to use, particularly the Python package <A href="https://help.sap.com/doc/cd94b08fe2e041c2ba778374572ddba9/2026_1_QRC/en-US/hana_ml.html" target="_blank" rel="noopener noreferrer">hana_ml</A>.</P><P>A popular approach for bringing in the latest documentation is by adding <A href="https://github.com/upstash/context7" target="_blank" rel="noopener nofollow noreferrer">Context7</A> to the project. There may be alternatives, but Context7 worked well for me. It can be added as MCP server both as local installation and through a remote URL. The local installation through the Cline interface didn't succeed for me, but the remote version worked very well. Funnily enough, you can just ask Cline to add the remote version itself.</P><P>Go into the "Plan" mode and request:</P><TABLE border="1" width="100%"><TBODY><TR><TD width="100%"><FONT color="#3366FF">Add an MCP server for Context7, use the remote URL</FONT></TD></TR></TBODY></TABLE><P>When happy with the plan, switch to the "Act" mode. Confirm the requests if you agree. And the MCP server for Context7 should have been added.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="090 context7.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418369i809A28DF0400E724/image-size/large?v=v2&amp;px=999" role="button" title="090 context7.png" alt="090 context7.png" /></span></P><P>&nbsp;</P><P>Verify this in the Cline settings. In the "MCP Servers" section under "Configure" the new MCP Server should show up. Also, below the prompt window you can see that the new MCP server is enabled and ready for use.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="100 context7 verification.gif" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418373i3D1F8104AD7D6994/image-size/large?v=v2&amp;px=999" role="button" title="100 context7 verification.gif" alt="100 context7 verification.gif" /></span></P><P>&nbsp;</P><H1 id="toc-hId-705548770">Vibe code HANA Machine Learning</H1><P>Everything we need has been set up. Let's do something useful with this it. This <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-machine-learning-with-sap-hana-cloud-and-bas/ba-p/14404096" target="_blank">blog</A> has a step-by-step explanation for writing the code to carry out a time-series forecast in SAP HANA Cloud. Hopefully that tutorial is useful to learn how the Machine Learning works, how Python code can instruct the Predictive Analysis Library to create the forecasts, without having to extract the data. Let's put vibe coding to the test, whether it can create such code for us.</P><P>First ensure that you have the credentials.json file in your folder, as shown in the same blog.</P><P>Then ask your code assistant in "Plan" mode to create the code.</P><TABLE border="1" width="100%"><TBODY><TR><TD width="100%"><FONT color="#3366FF">Create a new notebook called “HANA ML Time-series forecast”.&nbsp; Use the details from @/credentials.json to and the hana_ml package to connect to SAP HANA Cloud. Connect to the table HOTELNIGHTS. Do not download the data. Use hana_ml to have HANA Cloud aggregate this table on the MONTH column, summarising the column HOTELNIGHTS. Use the hana_ml Package to train an AdditiveModelForecast model and predict the next 6 months. Join the actuals and the prediction in a single structure and save as table HOTELNIGHTS_FORECAST_BAS. Test the code and fix any errors. Use context7.</FONT></TD></TR></TBODY></TABLE><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="110 timeseries prompt.png" style="width: 502px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418374iD34729B331E9447F/image-size/large?v=v2&amp;px=999" role="button" title="110 timeseries prompt.png" alt="110 timeseries prompt.png" /></span></P><P>&nbsp;</P><P>It is coming back with a very detailed, pretty looking plan.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="120 timeseries plan.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418375iA661CA92C2BDAC92/image-size/large?v=v2&amp;px=999" role="button" title="120 timeseries plan.png" alt="120 timeseries plan.png" /></span></P><P>&nbsp;</P><P>Click on "Act" to have it executed. Watch how Cline is going through that plan: writing, testing and potentially fixing the code. Pay attention to any Request prompts that might come up at the bottom left. Approve these if you agree. After a few minutes you should have a Notebook!</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="130 timeseries notebook.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418377iBC7AC884E71CCBDC/image-size/large?v=v2&amp;px=999" role="button" title="130 timeseries notebook.png" alt="130 timeseries notebook.png" /></span></P><P>&nbsp;</P><P>In the new Notebook, first choose your Python kernel on the top right. Then hit "Run All" to test the logic. It's looking good! But not perfect... The dates for the future months, that are predicted, are not the first day of that month.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="150 timeseries forecast v1.png" style="width: 781px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418378i8858F9B9F99CB1B7/image-size/large?v=v2&amp;px=999" role="button" title="150 timeseries forecast v1.png" alt="150 timeseries forecast v1.png" /></span></P><P>&nbsp;</P><P>Back to the prompt , and request a change. Here I am skipping the "Plan" mode and request this directly in the "Act" mode.</P><TABLE border="1" width="100%"><TBODY><TR><TD width="100%"><FONT color="#3366FF">Change the code so that the forecasted dates are always the first day of the month.</FONT></TD></TR></TBODY></TABLE><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="160 timeseries change.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418379i1E4A434B33E86467/image-size/large?v=v2&amp;px=999" role="button" title="160 timeseries change.png" alt="160 timeseries change.png" /></span></P><P>&nbsp;</P><P>And indeed, now the months are correctly dated! The LLMs came up with many different way to create the DataFrame with the dates that need to be predicted. <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/hands-on-tutorial-vibe-coding-sap-hana-cloud-machine-learning-with-bas/bc-p/14425212/highlight/true#M191435" target="_self">You can see a few in the comments</A>.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="170 timeseries forecast v2.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418380iC1A11E0B5143DCC1/image-size/large?v=v2&amp;px=999" role="button" title="170 timeseries forecast v2.png" alt="170 timeseries forecast v2.png" /></span></P><P>&nbsp;</P><P>As requested, the prediction is also saved to SAP HANA Cloud.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="180 timeseries forecast dbexplorer.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/418383iB2F2D47AC25F4854/image-size/large?v=v2&amp;px=999" role="button" title="180 timeseries forecast dbexplorer.png" alt="180 timeseries forecast dbexplorer.png" /></span></P><P>&nbsp;</P><H1 id="toc-hId-509035265">Summary</H1><P>It's impressive how efficient and helpful vibe coding can be!&nbsp;</P><P>It can be a massive time-saver, but do remember to verify the code...</P><P>&nbsp;</P><P>&nbsp;</P> 2026-06-18T11:12:53.472000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/going-beyond-the-tip-of-the-iceberg-with-sap-hana-cloud-sql-on-files/ba-p/14413871 Going Beyond the Tip of the Iceberg with SAP HANA Cloud SQL on Files 2026-06-24T04:34:53.984000+02:00 SeungjoonLee https://community.sap.com/t5/user/viewprofilepage/user-id/204092 <TABLE border="1" width="100%"><TBODY><TR><TD><STRONG>Related Blogs:<BR /></STRONG><UL><LI><A href="https://community.sap.com/t5/technology-blogs-by-sap/unlocking-the-true-potential-of-data-in-files-with-sap-hana-database-sql-on/ba-p/13861585" target="_blank">Unlocking the True Potential of Data in Files with SAP HANA Database SQL on Files in SAP HANA Cloud</A></LI><LI><A href="https://community.sap.com/t5/technology-blogs-by-sap/exploring-new-sql-on-files-use-cases-with-sap-hana-cloud-qrc-04-2024-and/ba-p/14076223" target="_self">Exploring New SQL on Files Use Cases with SAP HANA Cloud QRC 04/2024 and QRC 01/2025</A></LI><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/accelerating-your-analytics-with-delta-table-replication-into-sap-hana/ba-p/14290281" target="_self">Accelerating Your Analytics with Delta Table Replication into SAP HANA Cloud</A></LI></UL></TD></TR></TBODY></TABLE><P>As I mentioned in my previous blogs, SQL on Files, a capability of SAP HANA database in SAP HANA Cloud, has been continuously expanding since SAP HANA Cloud QRC 3/2024. We started with native read-only access to CSV, Parquet, and Delta tables in <A href="https://help.sap.com/docs/hana-cloud-data-lake/user-guide-for-data-lake-files/sap-hana-cloud-data-lake-administration-for-data-lake-files" target="_self" rel="noopener noreferrer">SAP HANA Cloud, data lake Files</A>. We then extended this access to external object storages (Amazon S3, Azure Blob/ADLS Gen2, Google Cloud Storage) and external Delta Sharing in QRC 4/2024. Most recently, we introduced Delta table-to-HANA replication in QRC 4/2025.</P><P>With SAP HANA Cloud QRC 2/2026, I'm pleased to announce that the journey continues, and this time the spotlight is on <STRONG>Apache Iceberg</STRONG>, the industry-standard open table format that has rapidly become a cornerstone of modern data lakehouse architectures. While Iceberg support has been quietly maturing in SAP HANA Cloud over the last few quarters, this is the first blog dedicated to telling the full story end-to-end.</P><P>In a nutshell, here is what SAP HANA Cloud now offers for Apache Iceberg, cumulatively from QRC 3/2025 through QRC 2/2026:</P><UL><LI><STRONG>Direct read-only access to Apache Iceberg tables</STRONG> in object storage via the <EM>file</EM> adapter, including time-travel queries (QRC 3/2025).</LI><LI><STRONG>Snapshot replica toggling support</STRONG> for Apache Iceberg virtual tables (QRC 1/2026).</LI><LI><STRONG>Data Page V2 (DATA_PAGE_V2) support</STRONG> in the Parquet reader, aligning SAP HANA Cloud with the latest Parquet ecosystem (QRC 1/2026).</LI><LI><STRONG>Direct read-only access to Apache Iceberg REST Catalogs</STRONG> via the new <EM>icebergcatalog</EM> adapter. QRC 1/2026 introduced the adapter with Cloudera and Databricks running on AWS as the initially certified providers, and QRC 2/2026 extends the multi-cloud coverage to Azure and GCP for Databricks while additionally certifying Snowflake across AWS, Azure, and GCP. Cloudera certification on Azure and GCP is planned for a future release.</LI><LI><STRONG>AliCloud Object Storage Service (OSS)</STRONG> added to the list of supported external object storages (QRC 2/2026).</LI></UL><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Overview.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/423444i566EB6374B5FACF2/image-size/large?v=v2&amp;px=999" role="button" title="Overview.png" alt="Overview.png" /></span></P><P>Alright, with these points in mind, let's dive deeper into the details with some examples.</P><P>&nbsp;</P><H3 id="toc-hId-1946346270">Apache Iceberg via the <EM>file</EM> Adapter</H3><P>Starting with SAP HANA Cloud QRC 3/2025, SQL on Files began supporting Apache Iceberg as an additional file format on top of the existing CSV, Parquet, and Delta table support. As highlighted in my <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/unlocking-the-true-potential-of-data-in-files-with-sap-hana-database-sql-on/ba-p/13861585" target="_self">first SQL on Files blog</A> and <A href="https://community.sap.com/t5/technology-blogs-by-sap/exploring-new-sql-on-files-use-cases-with-sap-hana-cloud-qrc-04-2024-and/ba-p/14076223" target="_self">follow-up blog</A>, the <EM>file</EM> adapter provides direct read-only access to files in <A href="https://help.sap.com/docs/hana-cloud-data-lake/user-guide-for-data-lake-files/sap-hana-cloud-data-lake-administration-for-data-lake-files" target="_self" rel="noopener noreferrer">SAP HANA Cloud, data lake Files</A> or in supported external object storages. With Iceberg now in the mix, customers can query Iceberg tables sitting on object storage without any data movement or ingestion.</P><P>Since Apache Iceberg is an open table format, the way of creating a virtual table is similar to Delta tables, where you only need to point the virtual table to the root of the Iceberg table:</P><pre class="lia-code-sample language-sql"><code>-- create a virtual table by pointing to an apache iceberg table CREATE VIRTUAL TABLE DEMO.ICEBERG_ORDERS AT "demo-hdlf_rs"."/iceberg/orders/" AS ICEBERG;</code></pre><P>A few things worth knowing about this support:</P><UL><LI>Both Iceberg <STRONG>v1 (Copy-on-Write)</STRONG> and <STRONG>v2 (Merge-on-Read)</STRONG> are supported for read-only access.</LI><LI>Time-travel queries are supported, similar to Delta tables, via either snapshot ID or timestamp:</LI></UL><pre class="lia-code-sample language-sql"><code>-- time travel by snapshot id SELECT * FROM DEMO.ICEBERG_ORDERS FOR VERSION AS OF '2953155114686890267'; -- time travel by timestamp SELECT * FROM DEMO.ICEBERG_ORDERS FOR SYSTEM_TIME AS OF '2025-04-08 07:43:31.793000000';</code></pre><UL><LI>Two new built-in functions help you explore Iceberg metadata directly from SAP HANA Cloud:</LI></UL><TABLE border="1" width="100%"><TBODY><TR><TD><STRONG>Function</STRONG></TD><TD><STRONG>Returns</STRONG></TD></TR><TR><TD width="50%">GET_ICEBERG_TABLE_SNAPSHOTS</TD><TD width="50%">Snapshot IDs, create time, sequence number, record count, file count, total file size</TD></TR><TR><TD width="50%">GET_VIRTUAL_TABLE_FILE_PARTITIONS</TD><TD width="50%">Partition values, file count, total file size per partition</TD></TR></TBODY></TABLE><UL><LI>The existing GET_REMOTE_SOURCE_FILE_COLUMNS built-in procedure has been extended to retrieve column information from Iceberg tables as well.</LI></UL><P><STRONG>Important:</STRONG> Apache Iceberg metadata stores absolute table paths. Simply copying Iceberg files from one location to another will cause inconsistencies, and creating virtual tables on the copied files will not work. Always create or migrate Iceberg tables using a proper Iceberg-compatible engine.</P><P>Please refer to the links below for further details.</P><UL><LI>SAP HANA Cloud, SAP HANA Database SQL on Files Guide: <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-sql-on-files-guide/create-virtual-table" target="_self" rel="noopener noreferrer">Create a Virtual Table</A></LI><LI>SAP HANA Cloud, SAP HANA Database SQL on Files Guide: <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-sql-on-files-guide/get-iceberg-table-snapshots" target="_self" rel="noopener noreferrer">GET_ICEBERG_TABLE_SNAPSHOTS Function</A></LI><LI>SAP HANA Cloud, SAP HANA Database SQL on Files Guide: <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-sql-on-files-guide/get-virtual-table-file-partitions" target="_self" rel="noopener noreferrer">GET_VIRTUAL_TABLE_FILE_PARTITIONS Function</A></LI><LI>SAP HANA Cloud, SAP HANA Database SQL on Files Guide: <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-sql-on-files-guide/get-remote-source-file-columns" target="_self" rel="noopener noreferrer">GET_REMOTE_SOURCE_FILE_COLUMNS</A></LI></UL><P>&nbsp;</P><H3 id="toc-hId-1749832765">Apache Iceberg REST Catalog via the <EM>icebergcatalog</EM> Adapter</H3><P>Reading Iceberg tables directly from object storage with the <EM>file</EM> adapter is great for self-managed scenarios. However, in real-world enterprise lakehouses, Iceberg tables are usually managed by an <STRONG>Iceberg REST Catalog</STRONG> server, for example, Databricks Unity Catalog, Cloudera Data Platform, or Snowflake. Customers told us that, for production-grade integration, the catalog-based approach is what they consider enterprise-ready: it centralizes table management, access control, and metadata governance.</P><P>That's where the new <EM>icebergcatalog</EM> adapter comes in. Starting with SAP HANA Cloud QRC 1/2026, SQL on Files supports direct read-only access to external Apache Iceberg REST Catalogs, with Cloudera and Databricks running on AWS as the initially certified providers. With QRC 2/2026, the multi-cloud coverage is extended to Azure (ADLS Gen2) and GCP (GCS) for Databricks, and Snowflake is added as an additionally certified provider across AWS, Azure, and GCP. Cloudera certification on Azure and GCP is planned for a future release.</P><P>The way to create a remote source to an Iceberg REST Catalog follows the official Apache Iceberg REST Catalog specification with bearer-token (OAuth) authentication:</P><pre class="lia-code-sample language-sql"><code>-- create a remote source to an iceberg rest catalog CREATE REMOTE SOURCE &lt;remote_source_name&gt; ADAPTER "icebergcatalog" CONFIGURATION ' provider=databricks; endpoint=&lt;rest_endpoint&gt;;' WITH CREDENTIAL TYPE 'OAUTH' USING 'access_token=&lt;bearer_token&gt;';</code></pre><P>Once the remote source is created, virtual tables are created in the same way as with other SAP HANA smart data access (a.k.a., SDA) adapters by pointing to the catalog → namespace → table:</P><pre class="lia-code-sample language-sql"><code>-- create a virtual table on a table exposed via the iceberg rest catalog CREATE VIRTUAL TABLE DEMO.VT_ORDERS AT "&lt;remote_source_name&gt;"."&lt;NULL&gt;"."&lt;namespace&gt;"."&lt;table&gt;";</code></pre><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Iceberg Catalog.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/423482i9CDBF8F4B7D99209/image-size/large?v=v2&amp;px=999" role="button" title="Iceberg Catalog.png" alt="Iceberg Catalog.png" /></span></P><P><SPAN>A few important notes on the certification scope and authentication:</SPAN></P><UL><LI><SPAN>The <EM>icebergcatalog</EM> adapter is designed to work with <STRONG>any spec-compliant Apache Iceberg REST Catalog</STRONG>. In theory, any REST catalog that follows the official spec and bearer-token authentication should work.</SPAN></LI><LI><SPAN>However, because small differences in authentication flows or vendor-specific extensions can cause unexpected issues, SAP officially certifies only a defined set of Iceberg REST Catalog providers per QRC, as summarized below. Other REST catalogs may work but are not officially certified.<BR /></SPAN></LI></UL><TABLE border="1" width="100%"><TBODY><TR><TD width="25%"><STRONG>Provider</STRONG></TD><TD width="25%"><STRONG>AWS</STRONG></TD><TD width="25%"><STRONG>Azure (ADLS Gen2)</STRONG></TD><TD width="25%"><STRONG>GCP (GCS)</STRONG></TD></TR><TR><TD width="25%">Cloudera</TD><TD width="25%">Certified (QRC 1/2026)</TD><TD width="25%">Planned</TD><TD width="25%">Planned</TD></TR><TR><TD width="25%">Databricks</TD><TD width="25%">Certified&nbsp;(QRC 1/2026)</TD><TD width="25%">Certified (QRC 2/2026)</TD><TD width="25%">Certified (QRC 2/2026)</TD></TR><TR><TD width="25%">Snowflake</TD><TD width="25%">Certified&nbsp;(QRC 2/2026)</TD><TD width="25%">Certified (QRC 2/2026)</TD><TD width="25%">Certified (QRC 2/2026)</TD></TR></TBODY></TABLE><UL><LI><SPAN>Authentication is currently limited to <STRONG>bearer-token (OAuth)</STRONG>. Full credential flows (e.g., automatic token acquisition and refresh by negotiating directly with the identity provider) are not implemented inside SAP HANA Cloud. Instead, the credential flow is expected to live in the application layer, while the database layer simply consumes the access token.</SPAN></LI></UL><P><SPAN>For renewing access tokens, the SET SESSION CREDENTIAL statement can be used to inject a freshly obtained access token from the application layer:</SPAN></P><pre class="lia-code-sample language-sql"><code>-- renew the access token from the application layer SET SESSION CREDENTIAL FOR REMOTE SOURCE &lt;remote_source_name&gt; TYPE 'OAUTH' USING 'access_token=&lt;access_token_string&gt;';</code></pre><P><SPAN>In other words, our expectation is that the credential flows are implemented in the application layer, while the database layer supports the way of renewing access tokens.</SPAN></P><P><SPAN>Please refer to the links below for further details.</SPAN></P><UL><LI><SPAN>SAP HANA Cloud, SAP HANA Database SQL on Files Guide: <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-sql-on-files-guide/create-apache-iceberg-catalog-remote-source" target="_self" rel="noopener noreferrer">Create an Apache Iceberg Catalog Remote Source</A></SPAN></LI></UL><P><SPAN>&nbsp;</SPAN></P><H3 id="toc-hId-1553319260"><SPAN>More Improvements Worth Highlighting</SPAN></H3><P><SPAN>Beyond the two main pillars above, several smaller but meaningful improvements landed alongside this Iceberg journey. They are bundled here for completeness.</SPAN></P><H4 id="toc-hId-1485888474"><SPAN>Snapshot replica toggling for Apache Iceberg (QRC 1/2026)</SPAN></H4><P><SPAN>For customers who want better query performance on Iceberg tables, snapshot replica toggling is now supported when the source is an Apache Iceberg table. As with other SDA-based virtual tables, you can toggle a virtual table to a snapshot replica by adding a snapshot replica with the ALTER VIRTUAL TABLE statement, then refresh or drop it as needed:<BR /></SPAN></P><pre class="lia-code-sample language-sql"><code>-- toggle to a snapshot replica ALTER VIRTUAL TABLE DEMO.VT_ORDERS ADD SHARED SNAPSHOT REPLICA; -- refresh the snapshot replica with the latest data from the source ALTER VIRTUAL TABLE DEMO.VT_ORDERS REFRESH SNAPSHOT REPLICA; -- delete the replica ALTER VIRTUAL TABLE DEMO.VT_ORDERS DROP REPLICA;</code></pre><P><SPAN>A couple of important constraints to keep in mind:</SPAN></P><UL><LI><SPAN>Snapshot replica toggling always creates a full snapshot of the data the virtual table is pointing to. This means it does not help if the target Iceberg table holds a vast volume of data that cannot be replicated in full into SAP HANA Cloud. In such cases, the recommended approach is to keep federation only, or wait for the upcoming chunk-based replication discussed in <STRONG>Looking Ahead</STRONG>.</SPAN></LI><LI><SPAN>Unlike Delta tables, real-time replication via toggling is not yet supported for Iceberg. Real-time replication on Iceberg will become available together with the upcoming interval-based CDC support (see <STRONG>Looking Ahead</STRONG>).</SPAN></LI></UL><H4 id="toc-hId-1289374969"><SPAN>Data Page V2 support in the Parquet reader (QRC 1/2026)</SPAN></H4><P><SPAN>The Parquet ecosystem has been moving toward <STRONG>Data Page V2</STRONG> (page_type=DATA_PAGE_V2), and several modern engines, now produce Parquet files using this newer page format. Starting with QRC 1/2026, the SAP HANA Cloud Parquet reader supports DATA_PAGE_V2, ensuring smooth interoperability with these engines. This is a quiet but important enabler for the full Iceberg story, especially for the Apache Iceberg REST Catalog scenario with Snowflake-managed Iceberg.</SPAN></P><H4 id="toc-hId-1092861464"><SPAN>AliCloud Object Storage Service (OSS) support (QRC 2/2026)</SPAN></H4><P><SPAN>The list of supported external object storages keeps growing. With QRC 2/2026, AliCloud Object Storage Service (OSS) joins Amazon S3, Azure Blob/ADLS Gen2, and Google Cloud Storage as a first-class storage option for SQL on Files.</SPAN></P><P><SPAN>However, please note one limitation specific to AliCloud deployments:</SPAN></P><UL><LI>SQL on Files queries accessing Google Cloud Storage on AliCloud Cloud deployments are not supported.</LI></UL><P><SPAN>Please refer to the links below for further details.</SPAN></P><UL><LI><SPAN>SAP HANA Cloud, SAP HANA Database SQL on Files Guide: <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-sql-on-files-guide/toggle-between-virtual-tables-and-replica-tables" target="_self" rel="noopener noreferrer">Toggle Between Virtual Tables and Replica Tables</A><BR /></SPAN></LI></UL><P><SPAN>&nbsp;</SPAN></P><H3 id="toc-hId-767265240"><SPAN>FAQs</SPAN></H3><P><STRONG>Q:</STRONG> Can I write to an Apache Iceberg table from SQL on Files?<BR /><STRONG>A:</STRONG> No. SQL on Files is read-only by design, both via the <EM>file</EM> adapter and the <EM>icebergcatalog</EM> adapter. INSERT, UPDATE, and DELETE on virtual tables pointing to Iceberg are not supported. Write operations should be performed by Iceberg-native engines (e.g., Spark, the catalog provider's compute) on the source side.</P><P><STRONG>Q:</STRONG> Which Iceberg REST Catalog providers are officially certified?<BR /><STRONG>A:</STRONG> Cloudera and Databricks running on AWS were initially certified with QRC 1/2026. With QRC 2/2026, Databricks is additionally certified on Azure and GCP, and Snowflake is newly certified across AWS, Azure, and GCP. Cloudera certification on Azure and GCP is planned for a future release. See the certification table in the <EM>icebergcatalog</EM> adapter section above for the full per-QRC, per-hyperscaler matrix. Other spec-compliant REST catalogs may work but are not officially certified.</P><P><STRONG>Q:</STRONG> What kind of authentication is supported for the <EM>icebergcatalog</EM> adapter?<BR /><STRONG>A:</STRONG> Bearer-token (OAuth) only. Full credential flows are expected to be handled in the application layer, with the access token being injected into SAP HANA Cloud either at remote source creation time or via SET SESSION CREDENTIAL for refresh.</P><P><STRONG>Q:</STRONG> Can I replicate an Apache Iceberg table into a SAP HANA table?<BR /><STRONG>A:</STRONG> Today, only snapshot replica toggling is supported for Iceberg. Full remote subscription-based replication (initial load + interval-based CDC), equivalent to the Delta table replication introduced in QRC 4/2025, is on the roadmap (see <STRONG>Looking Ahead</STRONG>).</P><P><STRONG>Q:</STRONG> I'm running Google Cloud Storage on AliCloud. Can I use SQL on Files to read data from there?<BR /><STRONG>A:</STRONG> No. SQL on Files queries accessing Google Cloud Storage are not supported on AliCloud deployments. Other supported object storages remain accessible.</P><P><STRONG>Q:</STRONG> Is there any plan to support Apache Iceberg format V3?<BR /><STRONG>A:</STRONG> The Iceberg ecosystem is evolving toward V3, and we are actively defining our adoption strategy. See <STRONG>Looking Ahead</STRONG> for the high-level direction.</P><P><STRONG>Q:</STRONG> I'm already accessing Apache Iceberg tables via the <EM>file</EM> adapter. Do I need to change anything with the introduction of the <EM>icebergcatalog</EM> adapter?<BR /><STRONG>A:</STRONG> No. The two adapters cover different scenarios and coexist. If you have been using the <EM>file</EM> adapter to access Iceberg tables directly on object storage, your existing setup keeps working as-is. The <EM>icebergcatalog</EM> adapter is an additional option for customers who manage their Iceberg tables through an Iceberg REST Catalog (Cloudera, Databricks, or Snowflake).</P><P>&nbsp;</P><H3 id="toc-hId-570751735">Looking Ahead</H3><P>While this blog focuses on what is generally available with QRC 2/2026, there are a couple of directions we are actively exploring for the longer term:</P><UL><LI><STRONG>Apache Iceberg table-to-HANA replication:</STRONG> Today, Delta table replication (QRC 4/2025) gives customers chunk-based or one-step initial load combined with user-driven or scheduled CDC. We see strong demand for the equivalent capability on Apache Iceberg via both the <EM>file</EM> adapter and the <EM>icebergcatalog</EM> adapter, including Data Replication UI integration. Release timing has not yet been finalized.</LI><LI><STRONG>Apache Iceberg format V3:</STRONG> V3 introduces extended types (e.g., variant, geometry, geography, nanosecond timestamps), default values, row lineage, and binary deletion vectors. We plan to selectively enable V3 capabilities as the broader ecosystem matures, so customers get predictable behavior with V3 tables. Release timing has not yet been finalized.</LI></UL><P>These items reflect long-term direction rather than committed deliveries, so please treat them as forward-looking signals rather than firm release plans.</P><P>&nbsp;</P><H3 id="toc-hId-374238230">Conclusion</H3><P>With SAP HANA Cloud QRC 2/2026, the Apache Iceberg story in SQL on Files is now complete in its first chapter: direct read-only access to Iceberg tables via the <EM>file</EM> adapter, enterprise-grade integration with Iceberg REST Catalogs via the new <EM>icebergcatalog</EM> adapter across AWS, Azure, and GCP, plus a set of supporting improvements like snapshot replica toggling, Data Page V2, and AliCloud OSS coverage.</P><P>This continues SAP HANA Cloud's commitment to open table formats, meeting customers where their data already lives, whether that's a Delta Lake, an Iceberg lakehouse, or a mixture of both. As we remain committed to innovation, stay tuned for upcoming updates that will continue to expand and enrich your SAP HANA Cloud experience, including richer Iceberg replication and evolving Iceberg V3 support.</P> 2026-06-24T04:34:53.984000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/automated-sql-plan-management-to-recover-from-degraded-performance/ba-p/14419577 Automated SQL Plan Management to recover from degraded performance 2026-06-24T13:33:10.872000+02:00 Taesuk https://community.sap.com/t5/user/viewprofilepage/user-id/202976 <P>You apply an upgrade over the weekend. Monday morning, someone reports that a critical reporting query is running much slower. You dig into the plan cache, compare before/after execution plans, find the optimizer chose a different join order, capture the old plan, pin it — and hours later, things are back to normal.</P><P>SAP HANA Cloud's <STRONG>SQL Plan Advisor</STRONG>, combined with <STRONG>SQL Plan Stability</STRONG>, now supports a fully automated detect-test-apply cycle that handles this class of degraded performance without manual intervention. Here is how it works and how to configure it.</P><H2 id="toc-hId-1817439420"><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="SPA.png" style="width: 400px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/425644iB43AB5BC4B71FA38/image-size/medium?v=v2&amp;px=400" role="button" title="SPA.png" alt="SPA.png" /></span></H2><H2 id="toc-hId-1620925915">The Two Pieces You Need to Know</H2><H3 id="toc-hId-1553495129">SQL Plan Stability — Your Insurance Policy</H3><P>SQL Plan Stability lets you capture an <STRONG>Abstract SQL Plan (ASP)</STRONG> for any SELECT statement. An ASP is a compact, engine-independent representation of an execution plan's key structural decisions: join order, join types, access paths, aggregation strategies. It is stored persistently and can be used later to <STRONG>regenerate</STRONG> that same physical plan, even after a software update or statistics refresh that would otherwise cause the optimizer to choose something different.</P><P><STRONG>Key points:</STRONG><BR />- SAP HANA Cloud now automatically collects ASP for every distinct HEX plans and corresponding statistics.<BR />- For non-HEX, only the statistics are collected for the corresponding statement hash</P><P>As ASPs are collected for every compiled plans with its corresponding statistics, we have a clearer understanding of what is a good plan versus a bad plan. This information is used it to protect against future degradation.</P><H3 id="toc-hId-1356981624">SQL Plan Advisor — The Automated Testing Engine</H3><P>SQL Plan Advisor is the active component. It continuously monitors SQL execution, detects performance degradation by comparing the current execution plans statistics over the statistics stored in SQL Plan Stability, and now&nbsp; <STRONG>automatically runs A/B tests between the current plan and candidate alternatives</STRONG>&nbsp;to find and validate a better one.</P><P>The candidate plan types it evaluates are:</P><TABLE border="1" width="100%"><TBODY><TR><TD width="50%">Candidate Type</TD><TD width="50%">What It Is</TD></TR><TR><TD width="50%">NON_HEX</TD><TD width="50%">A plan generated by the non-HEX execution engine, tested when the HEX plan degrades</TD></TR><TR><TD width="50%">CAPTURED_ABSTRACT_SQL_PLAN</TD><TD width="50%">An ASP previously captured via SQL Plan Stability</TD></TR><TR><TD width="50%">PLAN_VARIANT</TD><TD width="50%"><P>An alternative plan from the plan cache, typically tied to a different parameter binding by caching multiple plans for a given statement hash</P><P><STRONG>** NOT AUTOMATICALLY TESTED **</STRONG></P></TD></TR></TBODY></TABLE><P>&nbsp;</P><P>For parameterised queries, it is first determined if it can be covered by Plan Variant based on selectivity. If covered, the SQL Plan Advisor can test and collect statistics over the current plan and 50 with each alternative and checks for deviance before issuing a recommendation. <STRONG>As the testing maybe extensive, Plan Variant is currently scoped out for the automated testing</STRONG>.&nbsp;</P><P>&nbsp;</P><H2 id="toc-hId-1031385400">The Fully Automated Flow</H2><P>When configured for full automation, the pipeline looks like this:</P><P class="lia-align-center" style="text-align: center;"><BR />Query executes<BR />│<BR />▼<BR />HANA detects degradation (execution / CPU time compared to history)<BR />│<BR />▼<BR />Statement auto-registered in SQL Plan Advisor<BR />│<BR />▼<BR />Advisor tests current plan vs. candidates (NON_HEX, ASP)<BR />over multiple live executions — no downtime, no test window<BR />│<BR />▼<BR />Winner identified by empirical statistics<BR />│<BR />▼<BR />Better plan auto-applied on the next execution<BR />│<BR />▼<BR />Monitoring views updated — DBA can review at any time<BR /><BR /></P><P>The DBA is not in the critical path. You are informed after the fact, not paged during the incident.</P><P>&nbsp;</P><H2 id="toc-hId-834871895">Configuration</H2><H3 id="toc-hId-767441109">Step 1 — Enable SQL Plan Stability Capture for Critical Queries</H3><P>Do this <STRONG>before</STRONG> your next update. In SAP HANA Cloud Central, navigate to <STRONG>SQL Plan Stability</STRONG> and enable capture for your most sensitive statements. For most instances, this should be already enabled.</P><P>Or use SQL directly:</P><pre class="lia-code-sample language-sql"><code>-- Enable capture for a specific statement hash ALTER SYSTEM ALTER SQL PLAN CACHE ENTRY '&lt;statement_hash&gt;' SET PLAN STABILITY CAPTURE PLAN = TRUE;</code></pre><P>Captured ASPs land in `M_SQL_PLAN_CACHE_EXECUTION_PLAN_STATISTICS` and are referenced in the `CAPTURED_ABSTRACT_SQL_PLAN` column of `M_SQL_PLAN_ADVISOR`.</P><P><STRONG>Step 2 — Configure SQL Plan Advisor Parameters</STRONG></P><P>In SAP HANA Cloud Central, navigate to <STRONG>SQL Plan Advisor</STRONG> and enable <STRONG>Register Queries,</STRONG> and set <STRONG>Test Queries</STRONG> and <STRONG>Apply Advice</STRONG> to <STRONG>Automatically</STRONG>&nbsp;which should be the default configuration</P><P>&nbsp;</P><H2 id="toc-hId-441844885">Monitoring</H2><P>You can monitor all of this graphically in <STRONG>SAP HANA Cloud Central → SQL Plan Advisor</STRONG>, where the UI shows registered queries, test progress, performance comparisons, and applied recommendations in one view.</P><P>&nbsp;</P><H2 id="toc-hId-245331380">How SQL Plan Stability and SQL Plan Advisor Work Together</H2><P>The two features are complementary, not redundant:</P><P>- <STRONG>SQL Plan Stability</STRONG> is <STRONG>reactive and defensive</STRONG>: capture a known-good plan, apply it when something goes wrong. It requires you to have captured an ASP beforehand.</P><P>- <STRONG>SQL Plan Advisor</STRONG> is <STRONG>proactive and exploratory</STRONG>: it detects degradation automatically, tests multiple candidates (including your captured ASPs), and applies the empirically best one.</P><P>The combined strategy:</P><P>1. <STRONG>Before a QRC update</STRONG>: Capture ASPs for your critical statements via SQL Plan Stability which should be already captured.<BR />2. <STRONG>After the update</STRONG>: If the optimizer picks a worse plan for any of those statements, SQL Plan Advisor detects the degradation, tests the captured ASP against the new plan, and automatically applies the better one. Or, if a non-HEX plan is now converted to a HEX plan but the performance have degraded then reverts the execution back to non-HEX.<BR />3. <STRONG>Ongoing</STRONG>: Advisor continues monitoring for new degradation across the full workload, not just the statements you pre-selected.</P><P>You get both a safety net (captured ASPs) and an autonomous recovery mechanism (Advisor auto-apply).</P><P>&nbsp;</P><H2 id="toc-hId-48817875">What Gets Automated vs. What Still Needs DBA Judgement</H2><H4 id="toc-hId--387247287">Fully automated:</H4><P>- Detection of performance degradation<BR />- Registration of degraded statements as test candidates<BR />- A/B testing based on historical plans statistics.&nbsp;<BR />- Application of the most performant plan</P><P>&nbsp;</P><H4 id="toc-hId--583760792">Still requires DBA attention:</H4><P>- Parameterized queries which can be covered by Plan Variant should be manually tested if Plan Variant is applicable.---</P><H2 id="toc-hId--193468283">Conclusion</H2><P>The automated testing was first applied with QRC 01/2026 version and we have been closely monitoring on a daily basis. Based on our telemetry data, 155 plans have been corrected back to a more performant plan based on historically captured ASPs. We have noticed up to 20 times performance increases after the recovery. Most of the reasons of degraded performance was due to statistics changes in the underly tables or plan changes for parameterized queries which couldn't be covered by Plan Variant.</P><P>We are planning to improve the coverage of SQL Plan Advisor and improve the usability of our tooling for better manage SQL plans and improve user experiences.</P><H2 id="toc-hId--389981788">References</H2><P>- <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-administration-guide/sql-plan-advisor" target="_self" rel="noopener noreferrer">SQL Plan Advisor — SAP Help Portal</A><BR />- <A href="https://learning.sap.com/courses/deep-diving-into-sap-hana-sql-performance-fluctuation/automating-sql-plan-management-with-sql-plan-advisor" target="_self" rel="noopener noreferrer">Automating SQL Plan Management with SQL Plan Advisor — SAP Learning</A><BR />- <A href="https://learning.sap.com/courses/deep-diving-into-sap-hana-sql-performance-fluctuation/leveraging-sql-plan-stability-for-performance-optimization)" target="_self" rel="noopener noreferrer">Leveraging SQL Plan Stability for Performance Optimization — SAP Learning</A><BR />- <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/protect-from-performance-regression-with-sql-plan-stability/ba-p/13377751)" target="_self">Protect from Performance Regression with SQL Plan Stability — SAP Community Blog</A><BR />- <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-administration-guide/using-sql-plan-advisor-in-sap-hana-cloud-central" target="_self" rel="noopener noreferrer">Using SQL Plan Advisor in SAP HANA Cloud Central — SAP Help Portal</A><BR /><BR /></P> 2026-06-24T13:33:10.872000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/replicating-sap-s-4hana-cds-view-entities-into-sap-hana-cloud-a-practical/ba-p/14419559 Replicating SAP S/4HANA CDS View Entities into SAP HANA Cloud: A Practical Guide 2026-06-24T13:33:30.571000+02:00 Taesuk https://community.sap.com/t5/user/viewprofilepage/user-id/202976 <P>In analytics, speed is everything—yet pulling real-time data from a live SAP S/4HANA system remains a bottleneck for many organizations. When your reports rely on remote queries, they are instantly vulnerable to network latency and cross-system traffic. CDS view replication changes the game by moving the data to where the processing happens: caching it directly into SAP HANA Cloud and automating background syncs for instant reporting.</P><P>&nbsp;</P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="CDS_Replication.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/425604iB970917C7A25D5E6/image-size/large?v=v2&amp;px=999" role="button" title="CDS_Replication.png" alt="CDS_Replication.png" /></span></P><P>This post walks through everything you need to know to set up, operate, and monitor CDS view replication using the SAP HANA smart data access **abapodbc** adapter which is an extension to the already available federation as introduced by my colleague <A href="https://community.sap.com/t5/user/viewprofilepage/user-id/204092" target="_self">SeungjoonLee</A>&nbsp;and my previous blog.</P><P>&nbsp;</P><TABLE border="1" width="100%"><TBODY><TR><TD><STRONG>Related Blogs:</STRONG><BR /><UL><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/taking-data-federation-to-the-next-level-accessing-remote-abap-cds-view/ba-p/13635034" target="_blank">Taking Data Federation to the Next Level: Accessing Remote ABAP CDS View Entities in SAP HANA Cloud</A></LI><LI><A class="" href="https://community.sap.com/t5/technology-blog-posts-by-sap/easily-create-data-snapshots-from-abap-environment-s-cds-views-in-sap-hana/ba-p/13920746" target="_blank">Easily create data snapshots from ABAP environment's CDS Views in SAP HANA Cloud, SAP HANA Database</A></LI></UL></TD></TR></TBODY></TABLE><H2 id="toc-hId-1817439360">&nbsp;</H2><H2 id="toc-hId-1620925855">What Is CDS View Entity Replication?</H2><P>At its core, CDS view replication lets you maintain a local copy of an SAP S/4HANA CDS view as a native table inside SAP HANA Cloud. Once the initial snapshot is loaded, subsequent updates are applied using **Change Data Capture (CDC)** — only inserts, updates, and deletes are fetched and replayed, not the entire dataset.</P><P>The result is a target table in SAP HANA Cloud that stays in sync with your S/4HANA source, with a freshness you control.</P><H3 id="toc-hId-1553495069">Why This Matters</H3><P>- <STRONG>Performance</STRONG>: Local queries beat remote queries — no cross-system round trips at runtime.<BR />- <STRONG>Availability</STRONG>: Your replica stays accessible even if the S/4HANA system is under maintenance.<BR />- <STRONG>Flexibility</STRONG>: You can project specific columns or filter rows at the subscription level, so you only replicate what you need.<BR />- <STRONG>Integration</STRONG>: The replicated table is a plain SAP HANA table — you can join it with other local data, build calculation views on top of it, or expose it via BTP services.</P><P>&nbsp;</P><H2 id="toc-hId-1227898845">The Two Replication Phases</H2><P>Replication is split into two distinct phases:</P><H3 id="toc-hId-1160468059">Phase 1: Initial Load</H3><P>The initial load takes a full snapshot of the CDS view and copies it into the target table. You trigger this with the `QUEUE` command. The system handles chunking and parallelism internally — you just monitor progress.</P><P>One important caveat: the initial load does **not** check whether the target table is empty. If you're reloading after a reset, truncate the target table first.</P><H3 id="toc-hId-963954554">Phase 2: Delta Load</H3><P>Once the initial load completes, only changes (inserts, updates, deletes) are replicated. You have two options here:</P><P>- <STRONG>System-scheduled replication</STRONG>&nbsp;(`DISTRIBUTE`): The system automatically runs delta replication at a configurable interval. Set it once, let it run.<BR />- <STRONG>User-driven replication</STRONG>&nbsp;(`DELTA LOAD`): You explicitly trigger replication whenever you want. Useful for controlled refresh scenarios or testing.</P><P>&nbsp;</P><H2 id="toc-hId-638358330">Setting It Up:&nbsp;</H2><H3 id="toc-hId-570927544">Pre-requisite A — Setup Source System</H3><UL><LI><A href="https://help.sap.com/docs/abap-cloud/abap-integration-connectivity/exposing-sql-services-for-data-integration?version=s4_hana" target="_self" rel="noopener noreferrer">Enabling Access to ABAP-Managed Data for System-External Consumers</A></LI><LI><A href="https://help.sap.com/docs/abap-cloud/abap-integration-connectivity/inbound-data-integration-using-sql?version=s4_hana" target="_self" rel="noopener noreferrer">Accessing ABAP-Managed Data from System-External Consumers</A><UL><LI><A href="https://help.sap.com/docs/abap-cloud/abap-integration-connectivity/data-federation-using-sql-service-and-smart-data-access-for-sap-hana-cloud?version=s4_hana" target="_self" rel="noopener noreferrer">Accessing ABAP-Managed Data from<SPAN>&nbsp;</SPAN><SPAN class="">SAP HANA Cloud</SPAN><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><SPAN class="">SAP HANA</SPAN></A></LI></UL></LI></UL><H3 id="toc-hId-374414039">Pre-requisite B — Configure Cloud Connector to Source&nbsp;</H3><UL><LI><A class="" title="Configure Cloud Connector before connecting to on-premise sources and using them in various use cases. In the Cloud Connector administration, connect the SAP Datasphere subaccount to your Cloud Connector, add a mapping to each relevant source system in your network, and specify accessible resources for each source system." href="https://help.sap.com/docs/SAP_DATASPHERE/9f804b8efa8043539289f42f372c4862/f289920243a34127b0c8b13012a1a4b5.html?locale=en-US&amp;state=PRODUCTION&amp;version=cloud" target="_blank" rel="noopener noreferrer">Configure Cloud Connector</A></LI><LI><A class="" title="https://help.sap.com/docs/connectivity/sap-btp-connectivity-cf/configure-access-control-rfc" href="https://help.sap.com/docs/connectivity/sap-btp-connectivity-cf/configure-access-control-rfc" target="_blank" rel="noopener noreferrer">Configure Access Control (RFC)</A><SPAN>&nbsp;</SPAN>in the<SPAN>&nbsp;</SPAN>SAP BTP Connectivity<SPAN>&nbsp;</SPAN>documentation</LI></UL><H3 id="toc-hId-177900534">Step 1 — Create a Remote Source</H3><P>You need an abapodbc adapter remote source that connects SAP HANA Cloud to your S/4HANA system. This is your network tunnel between the two systems.</P><H3 id="toc-hId--93844340">Step 2 — Create the Virtual Table and Target Table</H3><P>Create a virtual table from the remote source (this is your read-through proxy to the CDS view), and create a local target table that will hold the replicated data.</P><H3 id="toc-hId--290357845">Step 3 — Create the Remote Subscription</H3><P>This is where the replication is defined. The `CREATE REMOTE SUBSCRIPTION` statement ties everything together:</P><pre class="lia-code-sample language-sql"><code>-- Simplest form: replicate everything CREATE REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" ON "SOURCE_CDS_VIEW" TARGET TABLE "TARGET_TABLE"; -- Column projection: only replicate specific fields CREATE REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" AS (SELECT "COL1", "COL2", "COL3" FROM "SOURCE_CDS_VIEW") TARGET TABLE "TARGET_TABLE"; -- Row filter: only replicate rows matching a condition CREATE REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" AS (SELECT * FROM "SOURCE_CDS_VIEW" WHERE "DEPARTMENT" = 'SALES') TARGET TABLE "TARGET_TABLE";</code></pre><P>The `WHERE` clause supports `=`, `&lt;&gt;`, `&lt;`, `&lt;=`, `&gt;`, `&gt;=`, `BETWEEN`, `IN`, `LIKE`, and logical operators `AND`/`OR`/`NOT`. A few syntax rules apply: field names must be double-quoted, string values single-quoted, and the right-hand side of a comparison must be a literal (no field-to-field comparisons).</P><H3 id="toc-hId--486871350">Step 4 — Perform the Initial Load</H3><pre class="lia-code-sample language-sql"><code>ALTER REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" QUEUE;</code></pre><P>Monitor progress in the `M_REMOTE_SUBSCRIPTION_INITIAL_LOADS` system view. The number of chunks can increase dynamically during the load as the system generates new ones.</P><H3 id="toc-hId--683384855">Step 5 — Enable Delta Replication</H3><P>Choose your mode:</P><pre class="lia-code-sample language-sql"><code>-- Scheduled: run every 15 minutes ALTER REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" DISTRIBUTE INTERVAL 900; -- On-demand: trigger manually ALTER REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" DELTA LOAD;</code></pre><P>&nbsp;</P><H3 id="toc-hId--879898360">Load Behaviors: Normal vs. Upsert</H3><P>When you create the subscription, you choose how DML operations are applied to the target table.</P><TABLE border="1" width="100%"><TBODY><TR><TD width="25%">Behavior</TD><TD width="25%">INSERT</TD><TD width="25%">UPDATE</TD><TD width="25%">DELETE</TD></TR><TR><TD width="25%"><STRONG>Normal</STRONG></TD><TD width="25%">Applied directly</TD><TD width="25%">Applied directly</TD><TD width="25%">Applied directly</TD></TR><TR><TD width="25%"><STRONG>Upsert</STRONG></TD><TD width="25%">Applied as UPSERT</TD><TD width="25%">Applied as UPSERT</TD><TD width="25%">Converted to UPDATE with `CHANGE_TYPE = 'D'`</TD></TR></TBODY></TABLE><P>The upsert mode is useful when you want to preserve a history of deletions or need idempotent replay. It requires three additional metadata columns in the target table as shown below:</P><pre class="lia-code-sample language-sql"><code>CREATE TABLE "TARGET_TABLE" ( "ID" INT, "NAME" NVARCHAR(100), "CHANGE_TYPE" NVARCHAR(1), -- 'A' = insert/update, 'D' = delete "CHANGE_TIME" TIMESTAMP, "CHANGE_SEQ" BIGINT -- optional: portion identifier ); CREATE REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" ON "SOURCE_CDS_VIEW" TARGET TABLE "TARGET_TABLE" CHANGE TYPE COLUMN "CHANGE_TYPE" CHANGE TIME COLUMN "CHANGE_TIME" CHANGE SEQUENCE COLUMN "CHANGE_SEQ" UPSERT;</code></pre><H2 id="toc-hId--783008858">&nbsp;</H2><H2 id="toc-hId--979522363">Managing Replication Day-to-Day</H2><H3 id="toc-hId--1469438875">Adjusting the Schedule</H3><P>You can update the replication interval without triggering an immediate run — useful for shifting frequency during off-peak hours:</P><pre class="lia-code-sample language-sql"><code>ALTER REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" INTERVAL 1800; -- 30 minutes</code></pre><P>To immediately trigger replication AND update the interval:</P><pre class="lia-code-sample language-sql"><code>ALTER REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" DISTRIBUTE INTERVAL 1800;</code></pre><H3 id="toc-hId--1665952380">Suspending During Maintenance</H3><P>If your S/4HANA system is going down for maintenance, suspend replication at the remote source level. All subscriptions under that source are paused; running jobs are allowed to finish but no new ones are scheduled.</P><pre class="lia-code-sample language-sql"><code>ALTER REMOTE SOURCE &lt;remote_source_name&gt; SUSPEND CAPTURE; -- When the system is back up: ALTER REMOTE SOURCE &lt;remote_source_name&gt; RESUME CAPTURE;</code></pre><P>Note that `SUSPEND CAPTURE` only affects system-scheduled jobs. Manual `DELTA LOAD` commands still work.</P><H3 id="toc-hId--1862465885">Cancelling a Running Job</H3><P>If a replication job is consuming too many resources or is stuck, you can cancel it mid-execution:</P><pre class="lia-code-sample language-sql"><code>-- Find the running job SELECT CONNECTION_ID FROM M_REMOTE_REPLICATION_JOBS WHERE STATUS = 'RUNNING' AND REMOTE_SUBSCRIPTION_NAME = 'ABAP_SUBSCRIPTION'; -- Cancel it ALTER SYSTEM CANCEL SESSION '&lt;connection_id&gt;';</code></pre><P>The system rolls back any partial changes, restoring the target table to its pre-job state. The job is automatically rescheduled for its next interval.</P><H3 id="toc-hId--1890795699">Resetting vs. Dropping</H3><P><STRONG>Reset</STRONG>&nbsp;returns the subscription to its initial state without removing it. Use this when you need to re-run the full initial load:</P><pre class="lia-code-sample language-sql"><code>ALTER REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" RESET; TRUNCATE TABLE "TARGET_TABLE"; ALTER REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION" QUEUE;</code></pre><P><STRONG>Drop</STRONG>&nbsp;permanently removes the subscription and cleans up resources on the ABAP source. The target table is retained.</P><pre class="lia-code-sample language-sql"><code>DROP REMOTE SUBSCRIPTION "ABAP_SUBSCRIPTION";</code></pre><H2 id="toc-hId--1793906197">&nbsp;</H2><H2 id="toc-hId--1990419702">Monitoring: The System Views You Need</H2><P>SAP HANA provides a set of dedicated system views for CDS view replication. Here's when to use each one:</P><TABLE border="1" width="100%"><TBODY><TR><TD width="50%">View</TD><TD width="50%">Use Case</TD></TR><TR><TD width="50%"><STRONG>REMOTE_SUBSCRIPTIONS</STRONG></TD><TD width="50%">Check subscription configuration, status, and the current replicated version (`EXTERNAL_VERSION_ID`)</TD></TR><TR><TD width="50%"><STRONG>M_REMOTE_SUBSCRIPTIONS</STRONG></TD><TD width="50%">Runtime health and status of active subscriptions</TD></TR><TR><TD width="50%"><STRONG>M_REMOTE_REPLICATION_JOBS</STRONG></TD><TD width="50%">Full lifecycle of scheduled jobs — PLANNED, RUNNING, COMPLETED, FAILED, CANCELLED</TD></TR><TR><TD width="50%"><STRONG>M_REMOTE_SUBSCRIPTION_INITIAL_LOADS</STRONG></TD><TD width="50%">Initial load progress by chunk: status, record count, execution time</TD></TR><TR><TD width="50%"><STRONG>M_REMOTE_SUBSCRIPTIONS_ABAPODBC</STRONG></TD><TD width="50%">ABAP-side replication ID and last processed portion ID</TD></TR></TBODY></TABLE><P>A few useful queries to keep handy:</P><pre class="lia-code-sample language-sql"><code>-- Is my initial load done? SELECT CHUNK_ID, TOTAL_RECORD_COUNT, EXECUTION_STATUS, EXECUTION_TIME FROM M_REMOTE_SUBSCRIPTION_INITIAL_LOADS WHERE REMOTE_SUBSCRIPTION_NAME = 'ABAP_SUBSCRIPTION'; -- What's my latest replicated version? SELECT EXTERNAL_VERSION_ID FROM REMOTE_SUBSCRIPTIONS WHERE SUBSCRIPTION_NAME = 'ABAP_SUBSCRIPTION'; -- Did any scheduled jobs fail recently? SELECT STATUS, ERROR_NUMBER, ERROR_MESSAGE, START_TIME, END_TIME FROM M_REMOTE_REPLICATION_JOBS WHERE REMOTE_SUBSCRIPTION_NAME = 'ABAP_SUBSCRIPTION' ORDER BY START_TIME DESC;</code></pre><H3 id="toc-hId-1814631082">&nbsp;</H3><H3 id="toc-hId-1618117577">Workload Management</H3><P>Automatic replication jobs participate in SAP HANA's native admission control. You can define workload class thresholds to prevent replication from overwhelming the system:</P><pre class="lia-code-sample language-sql"><code>CREATE WORKLOAD CLASS "AbapReplicationControl" SET 'ADMISSION CONTROL REJECT CPU THRESHOLD' = '90', 'ADMISSION CONTROL QUEUE CPU THRESHOLD' = '70'; CREATE WORKLOAD MAPPING "AbapReplicationMapping" WORKLOAD CLASS "AbapReplicationControl" SET 'APPLICATION NAME' = '_SYS_REMOTE_REPLICATION_ABAP_DELTA_LOAD';</code></pre><P>Jobs that exceed the QUEUE threshold are delayed; jobs that exceed the REJECT threshold fail with an error and are rescheduled.</P><H2 id="toc-hId-1715007079">&nbsp;</H2><H2 id="toc-hId-1518493574"><STRONG>Limitations to Know Before You Start</STRONG></H2><P>Before you commit to a schema design, check these constraints:</P><P>- <STRONG>Load behavior</STRONG>: Only Normal and Upsert are supported. INSERT and ARCHIVE load behaviors are not.&nbsp; (Archived data in the source are converted to deletes and will be removed from the target table)<BR />- <STRONG>Schema changes:</STRONG> The `WITH SCHEMA CHANGES` clause is not supported. If the CDS view schema changes, replication will fail and you'll need to manually update the target table and recreate the subscription.<BR />- <STRONG>Initial load:</STRONG> Only the automated `QUEUE` command is supported. The manual `INITIAL LOAD PREPARE/EXECUTE/FINALIZE` flow is not available.<BR />- <STRONG>Subquery expressions:</STRONG> Column projection and row filtering work. Function calls like `UPPER()` or arithmetic expressions in the subquery do not.<BR />- <STRONG>Subscription query:</STRONG> Once created, the subscription query cannot be changed. `ALTER REMOTE SUBSCRIPTION … REFRESH DEFINITION` is not supported. Recreate it if you need to change the filter or projection.<BR />- <STRONG>Target type</STRONG>: Only `TARGET TABLE` is supported. `TARGET PROCEDURE` and `TARGET TASK` are not.<BR />- **NULL handling**: Empty values on the ABAP side (empty strings, zeros) may arrive as `NULL` in the target table. Account for this in downstream logic.</P><P>---</P><H2 id="toc-hId-1321980069">Summary</H2><P>CDS view replication with the abapodbc adapter gives you a clean, CDC-based pipeline from SAP S/4HANA into SAP HANA Cloud. The lifecycle is straightforward: create the subscription, run the initial load with `QUEUE`, then choose between scheduled replication (`DISTRIBUTE`) or on-demand updates (`DELTA LOAD`). Built-in system views give you full visibility into job status, and native workload management ensures replication doesn't compete with your critical workloads.</P><P>The main things to keep in mind:</P><P>1. Truncate the target table before re-running the initial load — the system won't do it for you.<BR />2. Schema changes on the source CDS view require a manual rebuild of the subscription.<BR />3. Orphaned replications can silently drain ABAP resources — check with `DHADM` after any failed `CREATE REMOTE SUBSCRIPTION`.<BR />4. Failed scheduled jobs are **not** automatically rescheduled — you need to re-issue `DISTRIBUTE` after resolving the error.</P><P>With that, you have everything you need to get replication running and keep it healthy in production.</P> 2026-06-24T13:33:30.571000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/banking-sentinel-a-multi-agent-credit-risk-copilot-on-sap-btp-built-on-a/ba-p/14426538 Banking Sentinel: A Multi-Agent Credit-Risk Copilot on SAP BTP, Built on a Mixed AI Stack 2026-06-24T15:03:44.334000+02:00 Shahid https://community.sap.com/t5/user/viewprofilepage/user-id/15422 <P>&nbsp;</P><H2 id="the-problem-we-are-solving" id="toc-hId-1818273445">1. The Problem We Are Solving</H2><P>Every bank faces the same invisible risk: a customer looks safe in isolation but is quietly part of a web of connected entities (family trusts, guarantor networks, subsidiary companies) whose combined exposure is far beyond what any single credit file shows.</P><P>A borrower with a modest loan today can be the linchpin of a group with ten times the exposure. By the time a human analyst pieces it together, it may be too late.</P><P>The traditional tools are spreadsheets, batch reports, and credit scorecards that look at one dimension at a time. A scorecard says the debt-to-income ratio is 7.20 times, already a concern, already on file as a breach. Fine. But it does not tell you:</P><UL><LI>That the same customer’s income contract expires in 82 days</LI><LI>That when it does, their effective DTI rockets to 32.05 times, over 4x the regulator’s limit</LI><LI>That their loan is under-collateralized by AUD 620,000</LI><LI>That the connected-party graph has no relationship edge for this customer at all. A guarantor obligation only surfaces through a separate table, easy to miss if you only check the graph</LI><LI>That a shared guarantor connects this customer to a family trust, and that correctly scoping group exposure to only the loans within that group (not a guarantor’s unrelated obligations elsewhere in the portfolio) is its own real engineering problem, not just a data-fetching one</LI></UL><P>No single analyst. No single report. No single tool catches all of this at once.</P><P><STRONG>Banking Sentinel does.</STRONG></P><P><STRONG><span class="lia-unicode-emoji" title=":play_button:">▶️</span><A href="https://youtu.be/SajC-StscPo" target="_blank" rel="noopener nofollow noreferrer">Watch the demo on YouTube</A></STRONG></P><P><STRONG>GitHub:&nbsp;<A href="https://github.com/shahidla/Banking-Sentinel" target="_blank" rel="nofollow noopener noreferrer">https://github.com/shahidla/Banking-Sentinel</A></STRONG></P><P><STRONG><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="110.png" style="width: 400px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/425778i34A4B0687A131F03/image-size/medium?v=v2&amp;px=400" role="button" title="110.png" alt="110.png" /></span></STRONG></P><HR /><H2 id="why-this-matters" id="toc-hId-1621759940">2. Why This Matters</H2><P>A borrower’s risk doesn’t live in one place. It’s spread across payment behaviour, where their finances are heading, who they’re financially connected to, and whether there’s enough evidence to act on any of it. Miss one of those and a loss event can slip through a credit file that otherwise looks fine. Banking Sentinel checks all four, each step building on what the last one found, and turns the result into a single risk brief a human can review in under two minutes.</P><P>On the SAP side: this is a mixed-stack system with a deliberate boundary, not a claim that everything is SAP-native. The data and ML layer is SAP and proven (HANA Cloud, CAP, RPT-1, interchangeable with HANA PAL and Knowledge Graph Engine in production, see §11). The reasoning layer (the LLM, LangGraph, Langfuse) is not SAP, because AI Core isn’t available on the BTP trial tier this runs on. That’s an access constraint, not a verdict on SAP’s capability.</P><P>On the AI side: four patterns do real work here, not added for show. A ReAct loop drives graph traversal. A Reflexion-style critic (Reflection) checks evidence quality before anything is finalized. Human-in-the-Loop interrupts enforce compliance sign-off. A RAGAS-inspired claim-source check (cosine similarity, not the RAGAS library) catches an LLM asserting something the evidence doesn’t support, exactly what happens in §6 and §7.</P><HR /><H2 id="what-banking-sentinel-does" id="toc-hId-1425246435">3. What Banking Sentinel Does</H2><P>Banking Sentinel is a <STRONG>multi-agent AI risk intelligence system</STRONG> deployed on SAP BTP (Business Technology Platform). A risk analyst types a single sentence: <EM>“Analyse credit risk for customer 30100003.”</EM> Within roughly 75 to 110 seconds (HITL off or on), seven AI agents have examined the customer across four independent risk dimensions and produced an APRA-compliant risk brief, complete with regulatory references, confidence scores, and a clear recommendation.</P><P><STRONG>The system produces:</STRONG> - A risk score (0-100) with level: LOW / MEDIUM / HIGH / CRITICAL - Five specific findings with regulatory standard, severity, evidence source, and confidence score - Three actionable recommendations - Identified data gaps and uncertainties, not hidden, explicitly surfaced - A regulatory audit trail under APRA CPS 230 and APS 221 - A flag: APRA-Ready (true/false), whether the evidence is strong enough to take to a board</P><P><STRONG>The system refuses to:</STRONG> - Approve or reject a loan (it is a co-pilot, not a decision-maker) - Delete records - Override risk flags - Operate without a human sign-off when the risk is material</P><HR /><H2 id="the-technology-stack" id="toc-hId-1228732930">4. The Technology Stack</H2><P>Component Technology Purpose</P><TABLE><COLGROUP><COL /><COL /><COL /></COLGROUP><TBODY><TR><TD>Runtime Platform</TD><TD>SAP BTP Cloud Foundry</TD><TD>Deployment, scaling, service bindings</TD></TR><TR><TD>Application Framework</TD><TD>SAP CAP (CDS + Node.js)</TD><TD>OData APIs, HANA binding, CDS models</TD></TR><TR><TD>Primary Database</TD><TD>SAP HANA Cloud</TD><TD>All SAP TRBK/BCA banking tables</TD></TR><TR><TD>Vector Store</TD><TD>SAP HANA Cloud Vector Engine</TD><TD>APRA regulatory document embeddings</TD></TR><TR><TD>Graph Engine</TD><TD>GraphDB (RDF/SPARQL), HANA KGE in production</TD><TD>Connected-party traversal</TD></TR><TR><TD>Tabular AI Model</TD><TD>SAP RPT-1 (rpt.cloud.sap)</TD><TD>Tabular risk scoring without AI Core</TD></TR><TR><TD>Anomaly Detection</TD><TD>scikit-learn Isolation Forest (HANA PAL in production)</TD><TD>Statistical payment anomaly detection</TD></TR><TR><TD>LLM</TD><TD>Claude Haiku 4.5 (Anthropic)</TD><TD>All natural language reasoning</TD></TR><TR><TD>Agent Orchestration</TD><TD>LangGraph (StateGraph)</TD><TD>Multi-agent pipeline with conditional routing</TD></TR><TR><TD>State Persistence</TD><TD>PostgreSQL / Supabase</TD><TD>LangGraph checkpoint, survives CF restarts</TD></TR><TR><TD>Embeddings</TD><TD>OpenAI text-embedding-3-small</TD><TD>APRA document vectorisation</TD></TR><TR><TD>Observability</TD><TD>Langfuse</TD><TD>Per-agent token usage, latency, traces</TD></TR><TR><TD>Real-time UI</TD><TD>Server-Sent Events (SSE)</TD><TD>Live agent progress in browser</TD></TR><TR><TD>Event Mesh</TD><TD>Solace (Advanced Event Mesh)</TD><TD>Publishes pipeline events to <CODE>banking/*</CODE> topics. A real, persistent broker connection, genuinely sending messages, but no consumer is wired up within this demo; the UI’s live updates come via SSE, not Solace</TD></TR><TR><TD>Frontend</TD><TD>Vanilla HTML/CSS/JS</TD><TD>Bank-grade UI, no framework dependencies</TD></TR></TBODY></TABLE><H3 id="why-sap-rpt-1" id="toc-hId-1161302144">Why SAP RPT-1</H3><P>RPT-1 is SAP’s tabular foundation model, available via a public consumer API at rpt.cloud.sap without requiring AI Core or SAP AI Launchpad. It uses in-context learning: send it example rows from your portfolio with known risk categories, then ask it to classify your target customer. No training. No fine-tuning. Immediate results. Banking Sentinel uses it as the foundational risk score before any LLM reasoning begins.</P><H3 id="why-langgraph" id="toc-hId-964788639">Why LangGraph</H3><P>LangGraph is a graph-based agent orchestration framework. Each agent is a node. Data flows between nodes via a typed state object. Conditional edges allow the pipeline to branch: a low-risk customer skips the graph traversal and jumps straight to synthesis, while a high-risk customer goes through all four specialist agents. A re-query loop allows the Reflection agent to send the Relationship Agent back for a deeper traversal if the first pass was incomplete. This is impossible to express cleanly in a simple chain. It needs a state machine.</P><H3 id="why-hana-vector-engine" id="toc-hId-768275134">Why HANA Vector Engine</H3><P>APRA regulatory documents (APS 221, CPS 230, DTI Notices) are embedded into HANA Cloud’s native vector engine. When the Synthesis Agent writes the risk brief, it retrieves the most relevant regulatory clauses by semantic similarity, not keyword matching. This means the risk brief cites the actual paragraph of the actual regulation that applies to the specific risk being assessed.</P><HR /><H2 id="the-architecture" id="toc-hId-442678910">5. The Architecture</H2><PRE><CODE>User query → [Intake Agent] │ ┌─────────┼─────────┐ ▼ ▼ ▼ Simple Query Risk Rejection (direct Analysis (refuses DB lookup) pipeline) approvals) │ [Pattern Agent] RPT-1 + PAL + LLM │ Score &lt; 30 ──→ [Synthesis] Score ≥ 30 ──→ [Trajectory Agent] │ [Relationship Agent] ReAct loop: traverse graph, calculate exposure, check APRA │ [Reflection Check] LLM evaluates evidence quality │ Confidence &lt; 0.70 ──→ [Relationship] (re-query, max 2x) Confidence ≥ 0.70 OR max re-queries reached ──→ [Human Approval] │ ← interrupt ← Risk officer reviews and approves │ [Synthesis Agent] HANA Vector + APRA brief │ [END] Report persisted to HANA + AuditLog</CODE></PRE><P><STRONG>State flows through the pipeline in one typed object.</STRONG> Every agent reads everything every previous agent produced. The Synthesis Agent sees Pattern’s anomalies, Trajectory’s DTI projections, Relationship’s graph findings, and Reflection’s quality gaps, all at once. Nothing is lost between agents.</P><HR /><H2 id="the-seven-agents-a-full-walkthrough" id="toc-hId-246165405">6. The Seven Agents: A Full Walkthrough</H2><HR /><H3 id="agent-0-intake-agent" id="toc-hId-178734619">Agent 0: Intake Agent</H3><P><STRONG>Purpose:</STRONG> Parse the user’s query. Decide what kind of request this is. Route accordingly.</P><P><STRONG>The problem it solves:</STRONG> A risk officer might type “Analyse 30100001” or “What is the DTI for 30100001” or “Approve the loan for 30100001.” These are three completely different requests. The first triggers a full seven-agent pipeline. The second is a simple database lookup. The third must be refused. Banking Sentinel is a co-pilot, not a decision-maker.</P><P><STRONG>How it works:</STRONG> Claude Haiku reads the user’s query against a system prompt that defines three intent categories: - <CODE>RISK_ANALYSIS</CODE>: triggers the full agent pipeline - <CODE>SIMPLE_DATA_QUERY</CODE>: answered directly from HANA, no pipeline - <CODE>INAPPROPRIATE_REQUEST</CODE>: any attempt to approve, reject, delete, modify, or override</P><P>The agent also extracts the customer ID. SAP Business Partner numbers are 8-digit codes (e.g.&nbsp;30100001). The agent is instructed to extract them exactly, no reformatting.</P><P><STRONG>Output (stored in pipeline state):</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"isRiskAnalysis"</SPAN><SPAN class="">:</SPAN> <SPAN class="">true</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"isSimpleDataQuery"</SPAN><SPAN class="">:</SPAN> <SPAN class="">false</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"isInappropriateRequest"</SPAN><SPAN class="">:</SPAN> <SPAN class="">false</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"customerId"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"30100001"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"description"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"Analyse credit risk for customer 30100001"</SPAN></SPAN> <SPAN><SPAN class="">}</SPAN></SPAN></CODE></PRE></DIV><P><STRONG>What happens next:</STRONG> If <CODE>isRiskAnalysis</CODE>, the pipeline continues to the Pattern Agent. If <CODE>isSimpleDataQuery</CODE>, the query is delegated to a published MCP server (<CODE>cds-db-nlquery-mcp</CODE>) that translates it into a real CDS query, including JOINs and aggregates where needed, directly against <CODE>db/schema.cds</CODE>. A second Claude Haiku call then turns the result into a concise written answer. If <CODE>isInappropriateRequest</CODE>, a firm refusal is returned explaining that the system does not approve or override. That is a human decision.</P><P><STRONG>A live <CODE>SIMPLE_DATA_QUERY</CODE> example:</STRONG></P><P><EM>Query:</EM> <CODE>"What is the total loan amount across all customers"</CODE></P><P>This isn’t a single-row lookup. It requires a real SQL aggregate (<CODE>SUM(AMOUNT)</CODE>) across every loan in the portfolio, with no customer ID given at all. The Intake Agent correctly classifies it as <CODE>SIMPLE_DATA_QUERY</CODE> with <CODE>customerId: null</CODE>. The MCP server plans a structured query descriptor (<CODE>{"entity": "Loans", "aggregate": [{"fn": "sum", "col": "AMOUNT", "as": "total_loan_amount"}]}</CODE>), never raw SQL text, executes it as a real CDS query, and returns the result for Claude Haiku to phrase:</P><BLOCKQUOTE><P><STRONG>Total Loan Amount Across All Customers</STRONG> Total Loan Amount: AUD 31,773,000.00</P><P>Would you like a full risk analysis of any specific borrower</P></BLOCKQUOTE><P>Verified against live HANA: 30 loan records, AUD 31,773,000.00, confirmed against a direct <CODE>SUM(AMOUNT)</CODE> query as ground truth. This needed two rounds of fixes in the underlying <CODE>cds-db-nlquery-mcp</CODE> package: the query-planning LLM occasionally emitted a function-call string (<CODE>"SUM(AMOUNT)"</CODE>, then a wider variant with a trailing alias) as a literal column name instead of using the structured aggregate field. Both are now rejected up front with an actionable error instead of reaching HANA as a cryptic failure or, worse, silently returning a wrong total.</P><HR /><H3 id="agent-1-pattern-agent" id="toc-hId--93010255">Agent 1: Pattern Agent</H3><P><STRONG>Purpose:</STRONG> Establish the baseline risk signal. “Something feels wrong,” before any specific rule fires.</P><P><STRONG>The problem it solves:</STRONG> A credit scorecard tells you a number. But a number alone does not tell you whether the payment behaviour is suspicious, whether the debt structure is unusual, or whether there are statistical outliers in how this customer compares to the portfolio. Pattern Agent runs three methods and combines their signals.</P><P><STRONG>How it works:</STRONG> RPT-1 (Method 1) runs alone first. We diagnosed empirically that it was prone to BTP-only timeouts when racing the other two methods inside the same single-threaded event loop. Isolation Forest (Method 2) and the LLM narrative pass (Method 3) then run in parallel with each other, but not with RPT-1.</P><P><STRONG>Method 1: SAP RPT-1 (Tabular Foundation Model)</STRONG> The agent fetches up to 50 historical loan cases from <CODE>BCA_CREDIT_HISTORY</CODE>, a dedicated table of independently-labelled outcomes, to use as in-context examples. Each row has: case ID, DTI ratio, breach flag, total debt, annual income, and a known <CODE>arrears_outcome</CODE> (LOW/MEDIUM/HIGH/CRITICAL). It sends these to <CODE>rpt.cloud.sap/api/predict</CODE> alongside the target customer’s current profile (from <CODE>BCA_DTI</CODE>) with <CODE>arrears_outcome</CODE> marked <CODE>[PREDICT]</CODE>. RPT-1 applies in-context learning and returns a predicted arrears-risk category with a confidence score, mapped to a 0-100 scale via fixed floors (LOW:0, MEDIUM:26, HIGH:51, CRITICAL:76).</P><P><EM>For 30100003 (live run, 2026-06-24):</EM> RPT-1 returns <CODE>CRITICAL</CODE> with confidence 0.46.</P><P><STRONG>Method 2: Isolation Forest Anomaly Detection (scikit-learn / HANA PAL)</STRONG> The agent trains an Isolation Forest model on a 2D feature vector, payment delay days and dunning level (0-3), drawn from up to 500 rows combined from <CODE>DFKKOP</CODE> (open items) and <CODE>DFKKOPK</CODE> (cleared items), the bank’s full payment-history tables. Isolation Forest detects outliers by measuring how easily a data point can be isolated from the rest: anomalies are isolated in fewer splits. The 2D feature captures <EM>joint</EM> escalation: a customer whose delay <STRONG>and</STRONG> dunning level are both drifting gets flagged even when neither alone crosses a fixed threshold. Each of the customer’s payment rows is then scored: label <CODE>-1</CODE> for outlier, <CODE>1</CODE> for inlier.</P><P><EM>For 30100003:</EM> 0 of 7 payment rows flagged as outliers. This customer’s single loan has no statistically unusual payment behaviour on its own. (The risk signal here comes from RPT-1 and the LLM method below, not from this statistical check, a good illustration of why Pattern Agent runs three independent methods rather than relying on one.)</P><P><STRONG>Method 3: LLM Narrative Anomaly Detection (Claude Haiku)</STRONG> The raw customer data (loans, DTI record, up to 12 months of cleared payment history per loan, and collateral) is sent to Claude Haiku. The LLM is told the current APRA DTI threshold (fetched from <CODE>RegulatoryThresholds</CODE>, not hardcoded) and asked to find two specific things a single-row statistical check can’t: an <STRONG>escalating trend</STRONG> across a loan’s payment history (delay and/or dunning level rising over recent months, described narratively, e.g.&nbsp;“Loan L-001 deteriorated from on-time to 81-day delay / dunning level 3 over the last 5 months”), and <STRONG>under-collateralization</STRONG> (a loan’s amount exceeding the total value of its pledged collateral). It’s explicitly told not to flag a single row’s overdue days in isolation. That’s Method 2’s job.</P><P><EM>For 30100003:</EM> 2 anomalies identified, including loan L-004 (AUD 2.1M) exceeding its pledged collateral (AUD 1.48M) by AUD 620,000, a security deficiency.</P><P><STRONG>Tables read from HANA:</STRONG> - <CODE>bankingsentinel.BCA_CREDIT_HISTORY</CODE>: 50 labelled historical cases for RPT-1 in-context learning - <CODE>bankingsentinel.BCA_DTI</CODE>: this customer’s current DTI ratio, annual income, total debt, breach flag, income expiry. The row RPT-1 predicts against - <CODE>bankingsentinel.Loans</CODE>: loan IDs, amounts, types - <CODE>bankingsentinel.DFKKOP</CODE> / <CODE>bankingsentinel.DFKKOPK</CODE>: open and cleared payment transactions (delay days, dunning level) - <CODE>bankingsentinel.BCA_COLLATERAL</CODE>: collateral against loans - <CODE>bankingsentinel.RegulatoryThresholds</CODE>: current APRA DTI threshold</P><P><STRONG>Output (for 30100003):</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"riskScore"</SPAN><SPAN class="">:</SPAN> <SPAN class="">87</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"riskLevel"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"CRITICAL"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.46</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"signal"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"unclear"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"anomalies"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN></SPAN> <SPAN> <SPAN class="">"Loan L-004 (AUD 2.1M) exceeds collateral (AUD 1.48M) by AUD 620,000"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"..."</SPAN></SPAN> <SPAN> <SPAN class="">]</SPAN></SPAN> <SPAN><SPAN class="">}</SPAN></SPAN></CODE></PRE></DIV><P><STRONG>What happens next:</STRONG> If riskScore is below 30, the pipeline skips directly to Synthesis. No graph traversal needed. For 30100003 with score 87, the pipeline continues to the Trajectory Agent.</P><HR /><H3 id="agent-2-trajectory-agent" id="toc-hId--289523760">Agent 2: Trajectory Agent</H3><P><STRONG>Purpose:</STRONG> Project the customer’s financial position forward in time. “Where is this heading”</P><P><STRONG>The problem it solves:</STRONG> Current DTI is 7.20 times. The customer’s income is from a contract that expires in 82 days. When that contract ends, the customer has the same debt but a fraction of their current income remaining this year. What does DTI look like then And separately, even without the income expiry, what does a standard interest-rate-rise stress test do to this customer’s serviceability</P><P><STRONG>How it works:</STRONG></P><P><STRONG>Step 1: Fetch current DTI data from HANA</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">SELECT</SPAN> DTI_RATIO, TOTAL_DEBT, ANNUAL_INCOME, INCOME_EXPIRY, BREACH_FLAG</SPAN> <SPAN><SPAN class="">FROM</SPAN> bankingsentinel.BCA_DTI</SPAN> <SPAN><SPAN class="">WHERE</SPAN> PARTNER <SPAN class="">=</SPAN> <SPAN class="">'30100003'</SPAN></SPAN></CODE></PRE></DIV><P><EM>Result: DTI_RATIO=7.20, TOTAL_DEBT=2,100,000, ANNUAL_INCOME=291,667, INCOME_EXPIRY=‘2026-09-15’, BREACH_FLAG=true</EM></P><P><STRONG>Step 2: Fetch the APRA DTI threshold dynamically</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">SELECT</SPAN> LIMIT_PCT <SPAN class="">FROM</SPAN> bankingsentinel.RegulatoryThresholds</SPAN> <SPAN><SPAN class="">WHERE</SPAN> THRESHOLD_TYPE <SPAN class="">=</SPAN> <SPAN class="">'DEBT_TO_INCOME'</SPAN></SPAN></CODE></PRE></DIV><P>This is the key: the threshold is not hardcoded. It reads from the database, which can be updated in real time. For example, when the APRA Notice button is clicked in the UI, the threshold changes from 8.0 to 6.0 and the next run of the pipeline reflects the new regulatory position immediately.</P><P><STRONG>Step 3: Calculate Forward DTI (income-expiry projection)</STRONG></P><P>The formula: when income expires in N days, only N/365 of annual income remains effective this year.</P><PRE><CODE>Formula: effectiveIncome = annualIncome × (daysToExpiry / 365) futureDti = totalDebt / effectiveIncome For 30100003: daysToExpiry = 83 effectiveIncome = annualIncome × (83 / 365) ← only 23% of income remains futureDti = totalDebt / effectiveIncome = 32.05x</CODE></PRE><P><EM>Result: Forward DTI 32.05x, 301% above the APRA 8.0x limit.</EM></P><P><STRONG>Step 4: Calculate a second, independent projection, the rate-stress test</STRONG></P><P><CODE>trajectory-agent.js</CODE> runs a second forward calculation that has nothing to do with income expiry: APRA’s APG 223 serviceability buffer (a standard +3%, read dynamically from <CODE>RegulatoryThresholds.RATE_STRESS_BUFFER</CODE>) models what happens if the cost of servicing this debt rises uniformly. The debt side grows, income stays constant.</P><PRE><CODE>Formula: stressedDebt = totalDebt × (1 + RATE_STRESS_BUFFER_PCT / 100) futureDtiRateStress = stressedDebt / annualIncome For 30100003: stressedDebt = 2,100,000 × 1.03 = 2,163,000 futureDtiRateStress = 2,163,000 / 291,667 = 7.42x</CODE></PRE><P>This runs independently of Step 3. A customer could pass the income-expiry check but fail the rate-stress check, or vice versa. For 30100003 the income-expiry trajectory is already the dominant signal, but the rate-stress number (7.42x) is calculated and carried in the state regardless.</P><P><STRONG>Step 5: Check for conflicting signals</STRONG></P><P>The agent compares Pattern Agent findings against the DTI data to identify contradictions. For 30100003, four signals fired: 1. Income contract expires in 82 days. Primary servicing income at risk. 2. Active APRA DTI breach combined with imminent income loss. Compounding risk event. 3. Forward DTI of 32.0× projected. 301% above APRA limit post-expiry. 4. AUD 125,400 in scheduled loan payments fall within the income expiry window.</P><P><STRONG>Tables read from HANA:</STRONG> - <CODE>bankingsentinel.BCA_DTI</CODE>: DTI ratio, income, debt, expiry - <CODE>bankingsentinel.RegulatoryThresholds</CODE>: current APRA DTI threshold and the APG 223 rate-stress buffer - <CODE>bankingsentinel.Loans</CODE>: loan IDs for payment schedule lookup - <CODE>bankingsentinel.LoanSchedule</CODE>: scheduled payment amounts and due dates</P><P><STRONG>Output (for 30100003):</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"currentDti"</SPAN><SPAN class="">:</SPAN> <SPAN class="">7.2</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"futureDti"</SPAN><SPAN class="">:</SPAN> <SPAN class="">32.05</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"futureDtiRateStress"</SPAN><SPAN class="">:</SPAN> <SPAN class="">7.42</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"daysToExpiry"</SPAN><SPAN class="">:</SPAN> <SPAN class="">83</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"forwardPosition"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"DETERIORATING"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"conflictingSignals"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN></SPAN> <SPAN> <SPAN class="">"Income contract expires in 82 days — primary servicing income at risk"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Active APRA DTI breach combined with imminent income loss — compounding risk event"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Forward DTI of 32.0× projected — 301% above APRA limit post-expiry"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"AUD 125,400 in scheduled payments fall within income expiry window"</SPAN></SPAN> <SPAN> <SPAN class="">]</SPAN></SPAN> <SPAN><SPAN class="">}</SPAN></SPAN></CODE></PRE></DIV><P><STRONG>What happens next:</STRONG> Relationship Agent runs, receiving Pattern’s anomalies and Trajectory’s DTI context in its state.</P><HR /><H3 id="agent-3-relationship-agent" id="toc-hId--486037265">Agent 3: Relationship Agent</H3><P><STRONG>Purpose:</STRONG> Find every entity connected to this customer. Calculate their combined group exposure against APRA’s large exposure limit (APS 221).</P><P><STRONG>The problem it solves:</STRONG> Borrowers often operate within networks: family trusts, corporate groups, guarantor chains. APRA’s APS 221 standard requires banks to aggregate exposure across connected parties. A customer with a AUD 500,000 loan who is a member of a trust with AUD 2.5 million in loans is part of a AUD 3 million exposure group. The bank must know this.</P><P><STRONG>How it works, a ReAct (Reasoning + Acting) Loop:</STRONG></P><P>The agent uses Claude Haiku with three tools it can call iteratively:</P><P><STRONG>Tool 1: <CODE>hana_graph_traverse</CODE></STRONG> Traverses the graph database starting from the customer’s business partner ID. Returns all connected nodes (entities) and edges (relationship types) up to 8 hops deep. The underlying technology is GraphDB (an RDF triple store using SPARQL 1.1) in the demo environment, with HANA Knowledge Graph Engine as the production replacement. The same SPARQL queries work on both.</P><P>The data comes from <CODE>BUT050</CODE>, SAP’s Business Partner Relationship table, which stores every relationship between business partners: parent-subsidiary, guarantor, family trust member, director, and so on.</P><P><STRONG>Tool 2: <CODE>exposure_calculator</CODE></STRONG> Takes a list of all connected entity IDs and queries <CODE>BCA_GUARANTOR</CODE> and <CODE>Loans</CODE> to calculate total guaranteed and direct loan exposure across the group.</P><P><STRONG>Tool 3: <CODE>apra_threshold_check</CODE></STRONG> Checks the calculated exposure against APRA’s APS 221 large exposure limit (AUD 7.5 million for this demo). Returns the utilisation percentage.</P><P><STRONG>The ReAct loop in action, first pass:</STRONG></P><PRE><CODE>Step 1: LLM reasons → "I need to find connected parties for 30100003" LLM acts → calls hana_graph_traverse(startNode: "30100003", depth: 6) Tool returns → 1 enriched node, 0 chain edges (no BUT050 relationship row exists for this customer — the graph itself has no edge here) Step 2: LLM reasons → "Graph traversal found no connected entities. Check guarantor data directly." LLM acts → calls exposure_calculator(entityIds: ["30100003"]) Tool returns → groupExposure: 2,100,000 AUD (the guarantor cover on loan L-004, from BCA_GUARANTOR — found via the guarantor table, not via a graph edge) Step 3: LLM reasons → "Now check against APRA APS 221 threshold." LLM acts → calls apra_threshold_check(metricType: "aps221", value: 2100000, entityId: "30100003") Tool returns → aps221Pct: 28.0%, within_limit: true Step 4: LLM reasons → "Complete. I have enough to summarise." LLM stops calling tools and returns the finding.</CODE></PRE><P><STRONG>First-pass output:</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"nodes"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN><SPAN class="">"30100003"</SPAN><SPAN class="">]</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"groupExposure"</SPAN><SPAN class="">:</SPAN> <SPAN class="">2100000</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"aps221Pct"</SPAN><SPAN class="">:</SPAN> <SPAN class="">28.0</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.95</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"finding"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"Customer 30100003 and guarantor 30910005 (Rose Courtney) have combined APS 221 exposure of AUD 2.1M (28% of limit), with no regulatory breach."</SPAN></SPAN> <SPAN><SPAN class="">}</SPAN></SPAN></CODE></PRE></DIV><P>This first pass is real and grounded. It matches <CODE>BCA_GUARANTOR</CODE> exactly (Rose Courtney’s cover on L-004 is AUD 2.1M, the same as the loan amount). Notice what it <EM>missed</EM>: because <CODE>BUT050</CODE> has no relationship row for this customer, <CODE>hana_graph_traverse</CODE> alone returns zero edges. The only reason the guarantor showed up at all is that <CODE>exposure_calculator</CODE> separately queries <CODE>BCA_GUARANTOR</CODE> directly, a real gap between “what the graph shows” and “what the relational tables show.”</P><P><STRONG>A real bug we found and fixed while building this exact example.</STRONG> <CODE>exposure_calculator</CODE> originally summed a connected guarantor’s <EM>entire</EM> guarantee book, every loan they back anywhere in the portfolio, not just the loans within this customer’s group. Rose Courtney also guarantees loans for three other, unrelated customers elsewhere in the bank. Pulling those in inflated reported group exposure from a correct AUD 2.1M to a false AUD 11.78M (157% of the APS 221 limit) on a re-query, and it happened with two different starting customers who happen to share a guarantor, which is exactly the kind of bug that looks like a one-off coincidence until you check the data twice. The fix scopes guarantee cover to only loans already held by an entity in the connected group. The walkthrough below reflects the corrected behaviour, verified against live HANA after the fix.</P><P><STRONG>What the BUT050 table looks like</STRONG> (for a customer that <EM>does</EM> have graph relationships, e.g.&nbsp;30100001):</P><P>FROM_PARTNER TO_PARTNER RELTYP</P><TABLE><TBODY><TR><TD>30100001</TD><TD>30910005</TD><TD>CONTACT_PERSON</TD></TR><TR><TD>30100001</TD><TD>30910006</TD><TD>CONTACT_PERSON</TD></TR><TR><TD>30910005</TD><TD>30910006</TD><TD>FAMILY_TRUST_MEMBER</TD></TR></TBODY></TABLE><P><STRONG>Tables / databases accessed:</STRONG> - GraphDB / HANA KGE: SPARQL traversal of BUT050 relationship graph - <CODE>bankingsentinel.BCA_GUARANTOR</CODE>: guaranteed loan amounts - <CODE>bankingsentinel.Loans</CODE>: direct loan exposures - <CODE>bankingsentinel.BCA_SECTOR</CODE>: sector concentration check</P><P><STRONG>What happens next:</STRONG> Reflection Check evaluates whether this evidence is complete enough to proceed, and for 30100003, it doesn’t think it is.</P><HR /><H3 id="agent-4-reflection-check" id="toc-hId--682550770">Agent 4: Reflection Check</H3><P><STRONG>Purpose:</STRONG> The pipeline’s quality control layer. Evaluate its own work. Ask: “Do I have enough evidence to stand behind these findings”</P><P><STRONG>The problem it solves:</STRONG> The Relationship Agent found 1 node and 0 edges. But is that complete, or did the traversal stop because there was genuinely nothing more to find Is the CRITICAL Pattern Agent score consistent with a “clean” 28% exposure reading Are the anomalies actually linked to specific exposure items, or floating assertions without an evidence trail</P><P>A risk brief built on incomplete evidence is worse than no risk brief, because it creates false confidence.</P><P><STRONG>How it works:</STRONG></P><P>Reflection, a Reflexion-style critic step, means the LLM evaluates the prior agents’ outputs for quality rather than generating new findings. Claude Haiku is given a summary of all four agent findings and asked to assess four dimensions:</P><OL><LI><STRONG>Graph Completeness</STRONG>: Did the traversal stop early, or is there genuinely nothing more connected</LI><LI><STRONG>Signal Consistency</STRONG>: Do Pattern and Relationship findings agree CRITICAL risk score plus clean 28% exposure is an inconsistency worth questioning.</LI><LI><STRONG>Conflicting Signals</STRONG>: Are the trajectory conflicts explained by the graph, or still unresolved</LI><LI><STRONG>Evidence Trail</STRONG>: Is every risk claim backed by a specific TRBK record or exposure figure</LI></OL><P>The LLM returns a confidence score (0.0-1.0) and, if confidence is below 0.70, a specific <CODE>reQueryHint</CODE>, a targeted instruction for the Relationship Agent to go deeper.</P><P><STRONG>For 30100003, attempt 1 (live run, post-fix):</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"overallConfidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.58</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"gaps"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN></SPAN> <SPAN> <SPAN class="">"Graph traversal incomplete — only 1 node despite CRITICAL risk and a guarantor relationship claimed; guarantor 30910005 not traversed for counter-guarantees, parent entities, or cross-collateral exposure"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Pattern confidence (0.46) contradicts Relationship confidence (0.95) — unclear signal plus low pattern confidence suggests risk drivers not yet identified in the graph structure"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"APS 221 exposure (28%) appears manageable, but CRITICAL risk score (87) and deteriorating DTI (7.2 → 32.05) lack corresponding graph-based evidence; missing connection between income expiry and liability cascade"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"No TRBK defaults, arrears, or covenant breaches documented despite CRITICAL classification and imminent income loss"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Scheduled payment window (AUD 125,400 in 82 days) not mapped to specific loan tranches, guarantees, or maturity dates in the relationship graph"</SPAN></SPAN> <SPAN> <SPAN class="">]</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"reQueryHint"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"From customer 30100003, traverse all loan tranches and follow guarantor 30910005 for parent guarantees, cross-default clauses, and counter-guarantee relationships. Flag any loans with maturity at or under 82 days."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"reasoning"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"High trajectory risk is plausible but unmoored from relationship evidence; the single-node graph and low pattern confidence suggest the risk story exists operationally but hasn't been validated through full entity traversal."</SPAN></SPAN> <SPAN><SPAN class="">}</SPAN></SPAN></CODE></PRE></DIV><P><STRONG>Routing:</STRONG> 0.58 is below the 0.70 threshold, so it re-queries (attempt 1). Relationship Agent re-runs with the hint.</P><P><STRONG>Relationship Agent, re-query pass:</STRONG> This is where the bug above used to inflate the number. After the fix, the re-traversal genuinely adds one real fact, Rose Courtney’s family-trust connection to 30910006, without changing the exposure figure at all:</P><DIV class=""><PRE><CODE><SPAN><SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"nodes"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN><SPAN class="">"30100003"</SPAN><SPAN class="">]</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"groupExposure"</SPAN><SPAN class="">:</SPAN> <SPAN class="">2100000</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"aps221Pct"</SPAN><SPAN class="">:</SPAN> <SPAN class="">28.0</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.72</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"finding"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"Guarantor 30910005 (Rose Courtney) connected to family trust member 30910006; group exposure AUD 2.1M (28% of APS 221 limit) remains non-breaching. Loan tranche expiry, payment schedule, and cross-default clause details were not returned by traversal."</SPAN></SPAN> <SPAN><SPAN class="">}</SPAN></SPAN></CODE></PRE></DIV><P>Same AUD 2.1M. Same 28%. The only thing that changed is an honest new fact (the family-trust link) and an honest new gap (loan-level facility details aren’t in the graph at all). No fabricated crisis, because there’s no longer a code path that can produce one.</P><P><STRONG>For 30100003, attempt 2 (after the re-query):</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"overallConfidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.58</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"gaps"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN></SPAN> <SPAN> <SPAN class="">"Graph traversal halted at 1 node despite CRITICAL risk — no parent entities, upstream guarantors, or trust beneficiaries mapped beyond 30910005/30910006"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"APS 221 exposure at 28% (non-breaching) contradicts the CRITICAL risk score and 32x forward DTI — the exposure calculation may not reflect contingent liabilities or cross-default cascade"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Loan-level details (tranche expiry, payment schedule, cross-default triggers) explicitly noted as absent — Trajectory's 82-day income expiry and AUD 125,400 scheduled payments lack corresponding TRBK facility records"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Signal quality mismatch: rpt1Success and palSuccess both true with pattern confidence 0.46, but 'signal: unclear' prevents attributing CRITICAL to a specific driver"</SPAN></SPAN> <SPAN> <SPAN class="">]</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"reasoning"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"CRITICAL risk and a deteriorating forward DTI are unmoored from graph findings (1 node, non-breaching exposure); absent loan facility records and an incomplete guarantor/trust network prevent validating whether exposure cascade or cross-default mechanics justify the severity classification."</SPAN></SPAN> <SPAN><SPAN class="">}</SPAN></SPAN></CODE></PRE></DIV><P><STRONG>The routing decision:</STRONG> - Confidence stayed at 0.58, below 0.70, on <STRONG>both</STRONG> attempts, and the exposure figure stayed at the same correct AUD 2.1M on both attempts too - Maximum 2 re-queries reached. The pipeline does not loop forever. It proceeds anyway to Human Approval, carrying the gaps forward rather than blocking - The honest gap here isn’t “the LLM made something up.” It’s that this customer’s CRITICAL classification is genuinely driven by Pattern and Trajectory (the under-collateralized loan, the income-expiry DTI spike), not by the relationship graph, and the relationship graph’s own evidence (a clean, real 28%) can’t explain that on its own. Reflection correctly refuses to paper over that mismatch.</P><P><STRONG>Why this matters:</STRONG> Reflection is the only agent that looks at the entire pipeline’s output holistically. Here, it correctly distinguishes “this signal is genuinely clean” from “this signal explains the whole risk picture,” and won’t let a re-query manufacture false agreement between the two just to clear the confidence threshold.</P><HR /><H3 id="agent-5-human-in-the-loop-approval-hitl" id="toc-hId--879064275">Agent 5: Human-in-the-Loop Approval (HITL)</H3><P><STRONG>Purpose:</STRONG> Mandatory human checkpoint. No risk brief reaches the customer file without a human reviewing the evidence first.</P><P><STRONG>The problem it solves:</STRONG> APRA CPS 230 (Operational Resilience) requires that AI systems used in credit risk decisions operate as co-pilots, not autopilots. The risk officer must see the evidence and approve before the final brief is sealed.</P><P><STRONG>How it works:</STRONG></P><P>When Reflection says “proceed,” LangGraph halts the pipeline using <CODE>interruptBefore: ['humanApproval']</CODE>. The pipeline is paused. The state is persisted to PostgreSQL, which means the pause survives a server restart. The risk officer is notified in the UI that the pipeline is waiting for their review.</P><P>The risk officer sees: - Pattern findings: RPT-1 score, PAL anomaly counts, LLM anomalies - Trajectory: current DTI, forward DTI, days to income expiry - Relationship: the visual graph of connected entities with exposure amounts - Reflection gaps: explicitly what the AI was uncertain about</P><P>They click <STRONG>Approve</STRONG>. The pipeline resumes. Synthesis runs. The brief is sealed with <CODE>approvedBy: "risk.officer@bank.com.au"</CODE>.</P><P>If HITL is disabled (for demo or low-risk customers), the pipeline runs straight through to Synthesis automatically.</P><HR /><H3 id="agent-6-synthesis-agent" id="toc-hId--1075577780">Agent 6: Synthesis Agent</H3><P><STRONG>Purpose:</STRONG> Write the APRA-ready risk brief. Combine all four agents’ findings into a single, structured, regulatory-compliant document.</P><P><STRONG>The problem it solves:</STRONG> Four agents have run. Each produced findings in its own format. A risk officer needs one concise brief with clear findings, recommendations, regulatory citations, and an honest statement of what is uncertain. This brief must be good enough to take to a board. It must cite the actual APRA regulations that apply. It must acknowledge what is not yet known.</P><P><STRONG>How it works:</STRONG></P><P><STRONG>Step 1: Per-signal HANA Vector Search</STRONG> Instead of one generic regulatory query, the Synthesis Agent performs up to four targeted queries against the HANA Vector Engine, one per risk signal: - <EM>“DTI ratio 7.2 debt-to-income limit APRA activation”</EM> → retrieves APS 220 / DTI Notice clauses - <EM>“connected party group exposure APS 221 large exposure”</EM> → retrieves APS 221 thresholds - <EM>“income contract expiry forward DTI trajectory deteriorating”</EM> → retrieves forward assessment requirements - <EM>“CPS 230 operational resilience AI model governance audit trail”</EM> → retrieves CPS 230 obligations</P><P>The retrieved chunks are deduplicated and capped at 7 to stay within the token budget.</P><P><STRONG>Step 2: LLM Synthesis (Claude Haiku, maxTokens: 2500)</STRONG> All four agents’ outputs, plus the retrieved APRA regulatory text, are sent to Claude Haiku with a structured system prompt. The LLM produces a JSON risk brief. Findings are constrained to 20 words each (precision over prose), with one regulatory standard and one evidence source per finding.</P><P><STRONG>Step 3: Deterministic Guardrails</STRONG> The <CODE>apraReady</CODE> flag is NOT decided by the LLM. It is calculated deterministically from four conditions: confidence at or above 0.70, Reflection passed, regulatory docs retrieved, no regulatory context failure. This prevents the LLM from deciding its own work is ready.</P><P><STRONG>Step 4: Claim-Source Overlap Check (RAGAS-inspired, not the RAGAS library)</STRONG> A cosine-similarity check measures how much the LLM’s findings overlap with the retrieved regulatory text. Low overlap (below 30%) means the LLM may be relying on training data, or, as in this run, on an unverified claim from an earlier agent, rather than the retrieved documents. This is flagged in the uncertainty section.</P><P><STRONG>Step 5: Persist to HANA</STRONG> The risk assessment is written to <CODE>bankingsentinel.RiskAssessments</CODE>. Token counts and cost are written to <CODE>bankingsentinel.AuditLog</CODE>. Both are permanent records under CPS 230.</P><P><STRONG>For 30100003, the produced brief (live run, post-fix):</STRONG></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"riskScore"</SPAN><SPAN class="">:</SPAN> <SPAN class="">87</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"riskLevel"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"CRITICAL"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.46</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"findings"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN></SPAN> <SPAN> <SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"finding"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"Forward DTI 32.05x projected post-income expiry in 82 days; 301% above APRA limit."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"standard"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"DTI_NOTICE"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"severity"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"HIGH"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"evidenceSource"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"trajectory"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.46</SPAN></SPAN> <SPAN> <SPAN class="">}</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"finding"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"Loan L-004 AUD 2.1M exceeds collateral value AUD 1.48M by AUD 620K (LVR breach)."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"standard"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"APS221"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"severity"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"HIGH"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"evidenceSource"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"pattern"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.72</SPAN></SPAN> <SPAN> <SPAN class="">}</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"finding"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"Income contract expires in 82 days; primary servicing income at imminent risk."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"standard"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"DTI_NOTICE"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"severity"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"HIGH"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"evidenceSource"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"trajectory"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.46</SPAN></SPAN> <SPAN> <SPAN class="">}</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"finding"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"APS 221 exposure 28% non-breaching contradicts CRITICAL risk score; signal driver unclear."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"standard"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"APS221"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"severity"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"MEDIUM"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"evidenceSource"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"reflection"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.46</SPAN></SPAN> <SPAN> <SPAN class="">}</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">{</SPAN></SPAN> <SPAN> <SPAN class="">"finding"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"Graph traversal incomplete; guarantor network, cross-default triggers, loan facility records absent."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"standard"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"APS221"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"severity"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"MEDIUM"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"evidenceSource"</SPAN><SPAN class="">:</SPAN> <SPAN class="">"reflection"</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"confidence"</SPAN><SPAN class="">:</SPAN> <SPAN class="">0.58</SPAN></SPAN> <SPAN> <SPAN class="">}</SPAN></SPAN> <SPAN> <SPAN class="">]</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"recommendations"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN></SPAN> <SPAN> <SPAN class="">"Urgent loan-level facility query to validate tranche expiry, cross-default mechanics, and payment schedule alignment with the income expiry window."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Expand guarantor/trust network traversal to map contingent liabilities and cascade risk beyond the current 1-node graph."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Require customer income renewal evidence or alternative servicing capacity documentation before the 82-day expiry."</SPAN></SPAN> <SPAN> <SPAN class="">]</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"regulatoryRefs"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN><SPAN class="">"DTI_NOTICE"</SPAN><SPAN class="">,</SPAN> <SPAN class="">"APS221"</SPAN><SPAN class="">]</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"uncertainties"</SPAN><SPAN class="">:</SPAN> <SPAN class="">[</SPAN></SPAN> <SPAN> <SPAN class="">"Loan facility records missing; cannot validate the 82-day income expiry or AUD 125,400 scheduled payment timing."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Cross-default clause details absent; exposure cascade risk to guarantor 30910005 and family trust member 30910006 unquantified."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"Signal driver unresolved: rpt1Conf 0.46 with a CRITICAL classification is unmoored from APS 221 non-breach (28%) and 0/7 PAL anomalies."</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"CPS 230 guardrail: low claim-source overlap (10%) — findings warrant manual review"</SPAN></SPAN> <SPAN> <SPAN class="">]</SPAN><SPAN class="">,</SPAN></SPAN> <SPAN> <SPAN class="">"apraReady"</SPAN><SPAN class="">:</SPAN> <SPAN class="">false</SPAN></SPAN> <SPAN><SPAN class="">}</SPAN></SPAN></CODE></PRE></DIV><P>Notice what Finding 4 is actually saying: the APS 221 exposure genuinely doesn’t breach (28%, correct, verified), and that genuinely <EM>does</EM> sit next to a CRITICAL score, because the CRITICAL classification is driven by Trajectory and Pattern (the income-expiry DTI spike, the under-collateralized loan), not by group exposure at all. The brief doesn’t force these into false agreement. It states the real mismatch plainly, flags a 10% claim-source overlap, and sets <CODE>apraReady: false</CODE>. That’s the deterministic guardrail (Step 3) working as designed: a brief that’s honest about which of its own findings disagree, rather than one that resolves the disagreement by inventing a number nobody asked for.</P><HR /><H2 id="a-real-example-customer-30100003" id="toc-hId--978688278">7. A Real Example: Customer 30100003</H2><P>Let us trace the full pipeline from the moment the query is entered to the moment the risk brief appears on screen. This is a live run against the corrected <CODE>exposure_calculator</CODE> (see Agent 3), with HITL off (auto-advance through human approval).</P><P><STRONG>The query:</STRONG> <EM>“Analyse credit risk for customer 30100003”</EM></P><P><STRONG>Who is 30100003</STRONG> A retail customer with one business loan (L-004, AUD 2.1 million, retail property sector). Annual income AUD 291,667 from a contract expiring 2026-09-15, 82 days from this run. Current DTI is 7.20 times, already flagged in <CODE>BCA_DTI</CODE> as an active breach (<CODE>BREACH_FLAG: true</CODE>, dated 2025-08-10). On the surface this already looks concerning. What Banking Sentinel adds is <EM>how much worse</EM> it gets at income expiry, and, just as important, an honest account of which of its own signals do and don’t agree.</P><P><STRONG>The pipeline execution, phase sequence (real run, total 67.0s):</STRONG></P><P>Phase What happened</P><TABLE><COLGROUP><COL /><COL /></COLGROUP><TBODY><TR><TD>Intake</TD><TD>Classifies RISK_ANALYSIS, customerId=30100003</TD></TR><TR><TD>Pattern</TD><TD>RPT-1 → CRITICAL, confidence 0.46, score 87. Isolation Forest (scikit) → 0/7 outliers. LLM → 2 anomalies (L-004 under-collateralized by AUD 620K). Combined: score 87, CRITICAL, routes to high_risk.</TD></TR><TR><TD>Trajectory</TD><TD>Current DTI 7.20x. Forward DTI (income-expiry projection): 32.05x, 301% above the APRA 8x limit. Rate-stress DTI (independent +3% buffer check): 7.42x. Forward position: DETERIORATING. 4 conflicting signals raised.</TD></TR><TR><TD>Relationship (pass 1)</TD><TD>Graph traversal: 1 node, 0 edges (no BUT050 relationship row exists). Exposure calculator (via <CODE>BCA_GUARANTOR</CODE>, correctly scoped to this group): AUD 2.1M, 28% of APS 221 limit. Confidence 0.95.</TD></TR><TR><TD>Reflection (attempt 1)</TD><TD>Confidence 0.58, below the 0.70 threshold. 5 gaps raised, most pointed: a CRITICAL Pattern score next to a clean 28% exposure reading isn’t explained by the graph yet. Routes to re-query.</TD></TR><TR><TD>Relationship (pass 2, re-query)</TD><TD>Re-traversal still finds 0 new edges, but adds a real fact: guarantor 30910005 (Rose Courtney) is connected to family trust member 30910006. Exposure stays the same, correct AUD 2.1M, 28%. Confidence 0.72.</TD></TR><TR><TD>Reflection (attempt 2)</TD><TD>Confidence 0.58, unchanged. 4 gaps, all converging on the same honest point: the CRITICAL classification comes from Pattern and Trajectory, not from group exposure, and the relationship graph can’t be made to explain it. Maximum re-queries (2) reached, proceeds to Human Approval anyway, carrying the gaps forward.</TD></TR><TR><TD>Human Approval</TD><TD>HITL off for this run, auto-advances.</TD></TR><TR><TD>Synthesis</TD><TD>4 HANA Vector queries fire; 7 APRA regulatory chunks retrieved. Claude Haiku writes the brief. Claim-source overlap: 10% (low). Deterministic <CODE>apraReady</CODE> check: <STRONG>false</STRONG>.</TD></TR><TR><TD>Persist</TD><TD>Written to <CODE>RiskAssessments</CODE> and <CODE>AuditLog</CODE> in HANA. Risk brief delivered via SSE.</TD></TR></TBODY></TABLE><P><STRONG>Total pipeline time: 67.0 seconds (HITL off)</STRONG> <STRONG>Total cost: AUD 0.0025</STRONG> <STRONG>Tokens consumed: 2,919 input / 694 output</STRONG></P><P><STRONG>What the pipeline found that a scorecard would have missed:</STRONG> 1. The income contract expires in 82 days. Forward DTI rockets to 32.05x, 301% above the APRA limit. 2. Loan L-004 is under-collateralized by AUD 620,000, a security deficiency a single DTI number wouldn’t surface. 3. The relationship graph genuinely has no edge for this customer’s guarantor relationship. It only surfaces via a separate table, not the graph traversal tool. 4. The CRITICAL classification doesn’t come from group exposure at all, it comes from income expiry and under-collateralization, and the brief says so plainly rather than forcing every signal to agree with the headline score. 5. The risk brief explicitly says what it does not know: loan-level facility records, cross-default exposure to the guarantor and trust member, and why a CRITICAL score sits next to a clean exposure reading.</P><HR /><H2 id="regulatory-compliance-by-design" id="toc-hId--1175201783">8. Regulatory Compliance by Design</H2><P>Banking Sentinel is built to comply with three APRA standards. Every design decision traces back to a specific regulatory requirement.</P><H3 id="apra-aps-221-large-exposures" id="toc-hId--1665118295">APRA APS 221: Large Exposures</H3><P>APS 221 requires banks to aggregate exposure across connected parties and report when the total exceeds defined thresholds. The Relationship Agent exists solely to implement APS 221. Every graph traversal, every exposure calculation, every threshold check is an APS 221 obligation expressed in code.</P><H3 id="apra-cps-230-operational-resilience" id="toc-hId--1861631800">APRA CPS 230: Operational Resilience</H3><P>CPS 230 requires that AI systems used in risk decisions include human oversight, maintain audit trails, and survive operational disruptions. Three system design decisions implement this directly:</P><OL><LI><STRONG>Human-in-the-Loop interrupt</STRONG>: every material risk analysis pauses for human approval before the final brief is sealed</LI><LI><STRONG>PostgreSQL state persistence</STRONG>: the pipeline state survives CF restarts; an approval given before a deployment is not lost</LI><LI><STRONG>AuditLog</STRONG>: every pipeline run writes token counts, latency, model used, and cost to a permanent HANA table. The risk officer can reconstruct exactly what the AI did and why, months later.</LI></OL><H3 id="apra-dti-notice-debt-to-income-ratio" id="toc-hId--1889961614">APRA DTI Notice: Debt-to-Income Ratio</H3><P>The DTI Notice sets the threshold above which high-DTI lending requires additional oversight. Banking Sentinel reads this threshold dynamically from <CODE>RegulatoryThresholds</CODE>. When APRA changes its guidance, the threshold changes in the database. No code deployment required. The next pipeline run immediately reflects the new position.</P><P><STRONG>How the threshold actually changes: a real PDF upload, not a config flag.</STRONG> <CODE>POST /a2a/sync-apra</CODE> accepts a real APRA PDF (by URL or base64) and runs it through a genuine RAG ingestion pipeline:</P><OL><LI><STRONG>Extract</STRONG>: <CODE>pdf-parse</CODE> pulls the raw text out of the document.</LI><LI><STRONG>Chunk</STRONG>: an 800-character sliding window with 100-character overlap, so a sentence that spans a chunk boundary (e.g.&nbsp;“APS 221 requires…”) doesn’t lose its meaning at the seam.</LI><LI><STRONG>Embed</STRONG>: each chunk goes through OpenAI’s <CODE>text-embedding-3-small</CODE> (the same model Synthesis uses for retrieval, so the cosine-similarity comparison is apples-to-apples) and is stored in <CODE>bankingsentinel.RegulatoryDocuments</CODE>, a real HANA Vector Engine table, not an in-memory cache.</LI><LI><STRONG>Parse the actual new limit out of the text.</STRONG> For a <CODE>DTI_NOTICE</CODE> upload specifically, a small set of regex patterns (<CODE>DTI ≥ 6</CODE>, “DTI ratio greater than or equal to six times”, “debt-to-income … 6 times”, “DTI limit of 6”, and so on) extracts the real numeric threshold from the document’s own wording. It isn’t told the number in advance. If a number is found, <CODE>RegulatoryThresholds.LIMIT_PCT</CODE> is updated to that exact value; if parsing fails, it says so explicitly and leaves the threshold untouched rather than guessing.</LI><LI><STRONG>Replace, don’t accumulate</STRONG>: uploading a new <CODE>DTI_NOTICE</CODE> deletes the previous notice’s chunks first, so Synthesis’s vector search always retrieves the <EM>current</EM> regulatory position, not a mix of old and new guidance.</LI></OL><P>The risk officer’s experience: upload the new APRA notice once. The system reads its own new threshold out of the document, updates the live regulatory position, and replaces its own knowledge base. The very next <CODE>analyseRisk</CODE> call sees the new limit and the new source text, with zero deployment in between.</P><HR /><H2 id="what-the-risk-officer-sees" id="toc-hId--1793072112">9. What the Risk Officer Sees</H2><P>The Banking Sentinel UI displays the pipeline running in real time via Server-Sent Events (SSE). As each agent completes, its output appears on screen, bolded and populated.</P><P><STRONG>The dashboard shows</STRONG> (live values for the 30100003 run above): - <STRONG>Risk Score</STRONG>: 87 / 100, CRITICAL - <STRONG>Pattern Signal</STRONG>: unclear (2 anomalies) - <STRONG>RPT-1</STRONG>: CRITICAL, 46% confidence - <STRONG>Anomaly Detection (scikit-learn)</STRONG>: 0 / 7 payment rows flagged as outliers. This dashboard row shows whichever engine actually ran; it’s labelled “PAL” only when <CODE>ANOMALY_ENGINE=pal</CODE> is set and the HANA PAL service genuinely executed - <STRONG>Relationship Graph</STRONG>: interactive canvas; first pass 1 node / 0 edges, re-query pass still 1 node, adds the family-trust link to 30910006 without changing the exposure figure - <STRONG>Group Exposure</STRONG>: AUD 2,100,000, 28% of the APS 221 limit, consistent across both passes - <STRONG>Trajectory</STRONG>: current DTI 7.20x → forward DTI 32.05x (in 82 days); rate-stress DTI 7.42x - <STRONG>Reflection</STRONG>: confidence 0.58 (both attempts), 5 then 4 gaps, max re-queries reached - <STRONG>HITL Status</STRONG>: auto-approved (HITL off) for this run - <STRONG>Synthesis</STRONG>: full risk brief with findings, recommendations, regulatory refs, uncertainties; <CODE>apraReady: false</CODE> - <STRONG>Audit</STRONG>: cost AUD 0.0025, latency 67.0s, tokens 3,613 total (2,919 in / 694 out)</P><P>The report page merges all of this into a single printable brief. Every View Details panel shows the same data as the report, so what the risk officer approves is exactly what goes into the permanent record.</P><HR /><H2 id="key-design-decisions-and-lessons" id="toc-hId--1989585617">10. Key Design Decisions and Lessons</H2><H3 id="every-ai-call-has-a-named-pattern" id="toc-hId-1815465167">1. Every AI call has a named pattern</H3><P>There are no generic LLM calls. Every call is one of: intent classification (Intake), narrative anomaly detection (Pattern), quality self-evaluation (Reflection), ReAct tool-use loop (Relationship), or regulatory synthesis (Synthesis). When something breaks, you know exactly which pattern broke and why.</P><H3 id="the-apra-threshold-is-never-hardcoded" id="toc-hId-1618951662">2. The APRA threshold is never hardcoded</H3><P>Every agent that needs the DTI threshold reads it from <CODE>RegulatoryThresholds</CODE> at runtime. When the APRA Notice is applied in the UI, the threshold changes. The next run of the pipeline reflects it immediately. No deployment, no code change.</P><H3 id="langgraph-state-fields-must-be-declared" id="toc-hId-1422438157">3. LangGraph state fields must be declared</H3><P>LangGraph silently drops state fields that are not declared in <CODE>Annotation.Root</CODE>. This caused three bugs: <CODE>reflectionHistory</CODE>, <CODE>hitlEnabled</CODE>, and <CODE>totalLatencyMs</CODE> were all silently lost until each was explicitly declared with its reducer type.</P><H3 id="reflection-must-return-one-new-item-not-rebuild-the-full-history" id="toc-hId-1225924652">4. Reflection must return one new item, not rebuild the full history</H3><P>The <CODE>reflectionHistory</CODE> field uses an <CODE>append</CODE> reducer. Each time the node runs, its return value is appended to the existing array. If the node returns <CODE>[...existingHistory, newItem]</CODE>, the reducer appends the full rebuilt array again, producing duplicates. The fix: return only <CODE>[newItem]</CODE> and let the reducer do the accumulation.</P><H3 id="the-relationship-agent-must-not-re-traverse-from-arbitrary-nodes" id="toc-hId-1029411147">5. The Relationship Agent must not re-traverse from arbitrary nodes</H3><P>Early iterations of the Relationship Agent would, on finding 0 connections for the primary customer, attempt traversals from random other entities. This pulled in completely unrelated connected-party chains and inflated group exposure. The system prompt now explicitly instructs: if SPARQL returns 0 connections, that is expected. Use the guarantor data already returned.</P><H3 id="auditlog-and-state-persistence-are-separate-concerns" id="toc-hId-832897642">6. AuditLog and state persistence are separate concerns</H3><P>The <CODE>graph.updateState()</CODE> call persists data to the LangGraph checkpoint (PostgreSQL). The <CODE>logToAuditLog()</CODE> call writes to HANA. They are independent. If <CODE>graph.updateState()</CODE> throws and <CODE>logToAuditLog()</CODE> depends on it completing, the audit record is lost. The fix: wrap <CODE>graph.updateState()</CODE> in try-catch so <CODE>logToAuditLog()</CODE> always runs.</P><H3 id="a-connected-entitys-own-unrelated-obligations-are-not-this-groups-exposure" id="toc-hId-636384137">7. A connected entity’s own unrelated obligations are not this group’s exposure</H3><P><CODE>exposure_calculator</CODE> queried <CODE>BCA_GUARANTOR</CODE> filtered only on <CODE>GUARANTOR_PARTNER</CODE>, with no check on whose loan was actually being guaranteed. A guarantor connected to one customer’s group can also guarantee loans for entirely unrelated customers elsewhere in the portfolio, and that filter pulled in their <EM>entire</EM> guarantee book. This inflated reported APS 221 group exposure from a correct AUD 2.1M (28% of the limit) to a false AUD 11.78M (157%) on re-query, and it reproduced identically across two different starting customers who happened to share a guarantor, which is exactly the kind of bug that looks like a coincidence until you check the underlying data twice. The fix: a guarantee only counts toward this group’s exposure when the loan it covers is held by an entity already in the group, not merely guaranteed by one. Being connected to a customer’s group doesn’t mean every other loan that connection backs belongs to that group.</P><HR /><H2 id="what-comes-next" id="toc-hId-901457330">11. What Comes Next</H2><P>Banking Sentinel is a working prototype built on a SAP BTP trial account. The path to production involves four upgrades, each with a direct SAP equivalent:</P><P>Prototype Component Production Replacement</P><TABLE><COLGROUP><COL /><COL /></COLGROUP><TBODY><TR><TD>GraphDB sandbox (expires 7 days)</TD><TD>SAP HANA Knowledge Graph Engine</TD></TR><TR><TD>scikit-learn Flask service</TD><TD>SAP HANA PAL Isolation Forest (requires 3 vCPU)</TD></TR><TR><TD>Supabase free tier (pauses)</TD><TD>SAP BTP PostgreSQL Hyperscaler Option</TD></TR><TR><TD>Single demo customer (30100003)</TD><TD>Full portfolio: all BCA_DTI customers</TD></TR></TBODY></TABLE><P>The architecture does not change. The data sources do not change. The Isolation Forest model that runs on scikit-learn is the same algorithm as HANA PAL. The upgrade path is a configuration change, not a rebuild, but it’s worth being precise about what “the same queries work on both” actually means: the <EM>traversal semantics</EM> carry over, not literal byte-for-byte SQL. Here’s a concrete example for a sample business partner <CODE>0001</CODE>, showing what’s actually run today against GraphDB and the equivalent HANA KGE query it maps to:</P><P><STRONG>Today, SPARQL against GraphDB:</STRONG></P><PRE><CODE>PREFIX bs: &lt;urn:banking-sentinel:&gt; SELECT DISTINCT partnerId reltyp WHERE { &lt;urn:banking-sentinel:partner/0001&gt; bs:relatedTo* node . node bs:partnerId partnerId . OPTIONAL { &lt;urn:banking-sentinel:partner/0001&gt; rel node . BIND(STRAFTER(STR(rel), "relatedTo/") AS reltyp) } }</CODE></PRE><P><STRONG>Production target, HANA KGE, <CODE>GRAPH_TABLE</CODE> on a <CODE>BP_RELATIONSHIP_GRAPH</CODE> workspace</STRONG> (BUT050 rows as edges, BusinessPartners rows as vertices, the mapping described in <CODE>relationship-agent.js</CODE><span class="lia-unicode-emoji" title=":disappointed_face:">😞</span></P><DIV class=""><PRE><CODE><SPAN><SPAN class="">SELECT</SPAN> connected_partner_id, rel_type</SPAN> <SPAN><SPAN class="">FROM</SPAN> GRAPH_TABLE (BP_RELATIONSHIP_GRAPH</SPAN> <SPAN> MATCH (a<SPAN class="">:BusinessPartner</SPAN>)<SPAN class="">-</SPAN>[e<SPAN class="">:RELATED_TO</SPAN><SPAN class="">*</SPAN>]<SPAN class="">-&gt;</SPAN>(b<SPAN class="">:BusinessPartner</SPAN>)</SPAN> <SPAN> <SPAN class="">WHERE</SPAN> a.partner_id <SPAN class="">=</SPAN> <SPAN class="">'0001'</SPAN></SPAN> <SPAN> <SPAN class="">COLUMNS</SPAN> (</SPAN> <SPAN> b.partner_id <SPAN class="">AS</SPAN> connected_partner_id,</SPAN> <SPAN> e.rel_type <SPAN class="">AS</SPAN> rel_type</SPAN> <SPAN> )</SPAN> <SPAN>)</SPAN></CODE></PRE></DIV><P>Both express the same thing: variable-depth traversal from one business partner outward, returning every connected partner and the relationship type linking them. The vertex/edge mapping (BUT050 to graph edges, BusinessPartners to graph vertices) is the same in both. <STRONG>This KGE query is the documented target shape, not a verified one.</STRONG> HANA KGE isn’t available on the BTP trial tier this prototype runs on, so it has never actually executed against a live KGE instance. The honest claim is: the traversal logic and data model translate directly; the literal SQL above has not been run.</P><HR /><H2 id="summary" id="toc-hId-704943825">Summary</H2><P>Banking Sentinel is a complete, regulation-compliant, multi-agent AI risk system on a mixed stack with an explicit SAP boundary. SAP HANA Cloud, SAP CAP, and SAP RPT-1 carry the data and tabular-AI layer end to end, proven and ready to swap toward HANA PAL and HANA Knowledge Graph Engine in production. The reasoning and orchestration layer is non-SAP today because AI Core isn’t available on the trial tier this runs on, not because it was found wanting. Knowing exactly where that line sits is the more credible story.</P><P>Seven agents. Four risk dimensions. Three APRA standards. One risk officer decision. Under 80 seconds.</P><P>The customer who already shows a DTI breach on a scorecard, 7.20x, is actually a CRITICAL risk customer with a forward DTI of 32.05x at income expiry in 82 days, a loan under-collateralized by AUD 620,000, a connected-party graph with no relationship edge to show for it, and a CRITICAL classification that the brief is honest doesn’t come from group exposure at all, because group exposure (AUD 2.1M, 28%, stable across a re-query) genuinely doesn’t explain it.</P><P>Banking Sentinel does not approve or reject. It finds. It explains. It acknowledges what it does not know. And it gives the risk officer everything they need to make a confident, documented, APRA-compliant decision.</P><P>&nbsp;</P><P><EM>Built on: SAP BTP Cloud Foundry, SAP HANA Cloud, SAP CAP, SAP RPT-1, LangGraph, Claude Haiku 4.5, GraphDB / HANA KGE, scikit-learn / HANA PAL, Langfuse, Solace</EM></P><P><EM>APRA Standards: APS 221 (Large Exposures), CPS 230 (Operational Resilience), DTI Notice (Debt-to-Income)</EM></P> 2026-06-24T15:03:44.334000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/natural-language-queries-for-sap-cap-via-mcp-the-llm-translates-cds/ba-p/14426807 Natural Language Queries for SAP CAP via MCP: The LLM Translates, CDS Executes 2026-06-25T03:37:23.340000+02:00 Shahid https://community.sap.com/t5/user/viewprofilepage/user-id/15422 <P><STRONG>An MCP server that turns plain-English questions into CDS queries against your SAP CAP database layer, using schema annotations to resolve JOINs and disambiguate values.</STRONG></P><P data-unlink="true"><A href="https://github.com/shahidla/cds-db-nlquery-mcp" target="_blank" rel="noopener nofollow noreferrer">GitHub</A> · <A href="https://www.npmjs.com/package/@shahid.la/cds-db-nlquery-mcp" target="_blank" rel="noopener nofollow noreferrer">npm</A> · <A href="https://registry.modelcontextprotocol.io/?q=cds-db-nlquery-mcp" target="_self" rel="nofollow noopener noreferrer">MCP Registry&nbsp;</A></P><P>If you work with SAP CAP projects: a business question comes in, you open your SQL editor, figure out which tables hold the answer, and work out the JOINs and filter conditions. Ask a slightly different question and you start over.</P><P>This MCP server lets you ask the question directly instead.</P><PRE><CODE>Which customers have a DTI ratio above 5?</CODE></PRE><P>And get this back:</P><PRE><CODE>Results: 3 rows Partner : 30100003 DTI Ratio : 7.20 Customer : Domestic Customer AU 3 Partner : 30100001 DTI Ratio : 5.80 Customer : Domestic Customer AU 1 Partner : 30100004 DTI Ratio : 5.40 Customer : Domestic Customer AU 4</CODE></PRE><P>That's what <STRONG><A href="https://github.com/shahidla/cds-db-nlquery-mcp" target="_blank" rel="noopener nofollow noreferrer">cds-db-nlquery-mcp</A></STRONG> does. It's an MCP server that sits on top of your SAP CAP project and lets you query your database-layer entities (the ones defined in <CODE>db/schema.cds</CODE>, not your OData service layer) in natural language, from Claude Code, Claude Desktop, or any MCP-compatible host.</P><P>Here's how it works, why it's architected this way, and three real examples against a live HANA Cloud schema.</P><HR /><H2 id="toc-hId-1818276234">How It Works, At a Glance</H2><PRE><CODE> Your question (plain English) │ ▼ ┌────────────────────────┐ │ Stage 1: LLM │ reads your CDS schema (labels, associations, │ Question → Descriptor │ @Common.Text, enums) → outputs JSON only: └─────────┬───────────────┘ { entity, select, where, valCol, limit } │ JSON descriptor ▼ ┌────────────────────────┐ │ Stage 2: CDS │ descriptor → CQN → real SQL JOINs, │ Descriptor → SQL │ generated from association metadata, └─────────┬───────────────┘ not guessed by the LLM │ rows from HANA ▼ ┌────────────────────────┐ │ Stage 3: Answer │ MCP client renders the rows, or hands them │ Rows → Plain English │ to a second LLM call to phrase the answer └────────────────────────┘ The LLM never touches SQL. CDS never touches your question.</CODE></PRE><HR /><H2 id="toc-hId-1621762729">Why This Isn't "LLM Writes SQL"</H2><P>One common pattern for natural-language-to-SQL works by dumping the whole database schema into a prompt, asking the LLM to write SQL, and running whatever comes back. The LLM has to get table names, JOIN syntax, and HANA-specific SQL right in one shot, with no framework checking its work.</P><P>This package splits the problem in two:</P><OL><LI><STRONG>The LLM's only job is translation.</STRONG> It reads your CDS schema and turns a question into a small JSON descriptor: entity, columns, filters. It never writes SQL.</LI><LI><STRONG>The CDS framework's only job is execution.</STRONG> It turns that descriptor into a CQN query (CDS's own query representation) using association-path expressions, and lets CDS, not the LLM, not hand-written SQL, generate the actual SQL JOINs that HANA executes.</LI></OL><P>That division matters: two common failure modes are doing JOINs in JavaScript after fetching too much data, or trusting the LLM to write syntactically valid HANA SQL from scratch. This server avoids both. There's no JavaScript-side join, no post-fetch filtering. <CODE>WHERE</CODE>, <CODE>ORDER BY</CODE>, and <CODE>LIMIT</CODE> are all pushed down to HANA in a single round-trip per question.</P><HR /><H2 id="toc-hId-1425249224">Three-Stage Architecture</H2><PRE><CODE>You: "Show me active loans for customers in the mining sector, with the borrower's name and loan amount"</CODE></PRE><H3 id="toc-hId-1357818438">Stage 1: Question → JSON Descriptor</H3><P>At startup, the server loads your full CDS model and builds a compact schema description: entity names, columns with types, labels, enums, <CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1648847">@Common</a>.Text</CODE> references, and association paths. That schema text plus your question goes to a fast, cheap model tier (<CODE>claude-haiku-4-5-20251001</CODE>, <CODE>gpt-4o-mini</CODE>, or any configured provider). The LLM responds with a descriptor, not SQL:</P><PRE><CODE>{ "entity": "Loans", "select": ["LOAN_ID", "AMOUNT", "customer.BU_SORT1", "status.TEXT"], "where": [ { "col": "customer.SECTOR_CODE", "op": "=", "val": "MINING" }, { "col": "status.TEXT", "op": "like", "val": "Active" } ], "limit": 50 }</CODE></PRE><P>This is a cheap LLM call, a structured-output translation task, not a reasoning task. The model isn't answering your question; it's mapping it onto the schema vocabulary it was given.</P><H3 id="toc-hId-1161304933">Stage 2: Descriptor → CQN → SQL (CDS Generates the JOINs, Not the LLM)</H3><P>The server takes that descriptor and builds <STRONG>one</STRONG> CDS CQN query. <CODE>customer.BU_SORT1</CODE> and <CODE>status.TEXT</CODE> are CDS association paths. The server resolves them against the schema's join metadata and CDS compiles them into real SQL JOINs:</P><PRE><CODE>SELECT L.LOAN_ID, L.AMOUNT, BP.BU_SORT1 AS customer_BU_SORT1, LSC.TEXT AS status_TEXT FROM bankingsentinel_Loans L INNER JOIN bankingsentinel_BusinessPartners BP ON BP.PARTNER = L.PARTNER LEFT JOIN bankingsentinel_LoanStatusCodes LSC ON LSC.CODE = L.STATUS WHERE UPPER(BP.SECTOR_CODE) = 'MINING' AND UPPER(LSC.TEXT) LIKE '%ACTIVE%' LIMIT 50</CODE></PRE><P>The JOINs are real, generated by CDS from association metadata you already declared in your schema, not reconstructed from scratch by an LLM guessing at foreign keys.</P><H3 id="toc-hId-964791428">Stage 3: Results → Answer</H3><P>The rows come back from <CODE>cds.run()</CODE> as plain objects. Your MCP client (Claude Code, Claude Desktop, etc.) renders them, or, as in the Banking Sentinel example later, hands them to a second, separate LLM call whose only job is to phrase the answer in plain English.</P><HR /><H2 id="toc-hId-639195204">Wait, Is That Safe?</H2><P>Right question. The server only ever issues <CODE>SELECT</CODE>. There is no code path that builds an INSERT, UPDATE, or DELETE. Beyond that, there are three independent layers, all enforced server-side, none of them optional:</P><OL><LI><STRONG>Database-level</STRONG>: point <CODE>MCP_DB_USER</CODE>/<CODE>MCP_DB_PASSWORD</CODE> at a dedicated read-only HANA user. If the framework had a bug, HANA itself would still reject a write.</LI><LI><STRONG>Entity allowlist</STRONG> (<CODE>MCP_ALLOWED_ENTITIES</CODE><span class="lia-unicode-emoji" title=":disappointed_face:">😞</span> restricts which entities are queryable at all. This is enforced on the entity you asked for <STRONG>and</STRONG> on every entity reached through an association path, so allowlisting <CODE>BCA_DTI</CODE> but not <CODE>BusinessPartners</CODE> blocks <CODE>customer.BU_SORT1</CODE> from being selected at all, closing the obvious bypass-via-JOIN hole.</LI><LI><STRONG>Column blocklist</STRONG> (<CODE>MCP_BLOCKED_COLUMNS</CODE><span class="lia-unicode-emoji" title=":disappointed_face:">😞</span> strips named columns (e.g. <CODE>PASSWORD</CODE>, <CODE>EMBEDDING</CODE>, <CODE>SSN</CODE>) before the query is even built. They're never sent to HANA, not fetched-then-redacted.</LI></OL><P>The one thing worth saying plainly, because the README says it plainly: <STRONG>this bypasses your CAP service-layer <CODE>@requires</CODE>/<CODE>@restrict</CODE> annotations.</STRONG> It queries <CODE>db/schema.cds</CODE> entities directly via <CODE>cds.run()</CODE>, not your OData service. If your service layer is where your authorization model lives, that model doesn't apply here. The three controls above are the substitute, not an addition.</P><P>For a quick demo, start with none of them and everything is queryable. For anything pointed at real data, set all three.</P><HR /><H2 id="toc-hId-442681699">The Real Lever: Schema Annotations, Not Prompt Engineering</H2><P>Before the examples, this is the part that actually determines whether your queries come back right: <STRONG><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1790377">@nlp</a>.label</CODE></STRONG>.</P><P>The server already reuses your existing <CODE>@title</CODE> annotations automatically. If your schema already has Fiori value-help labels, the LLM gets those for free. <CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1790377">@nlp</a>.label</CODE> exists for the cases where <CODE>@title</CODE> doesn't fit or isn't disambiguating enough, and it overrides <CODE>@title</CODE> when both are present.</P><P>Here's a real one from the schema you'll see in the examples below:</P><PRE><CODE>@title: 'Partner Type Code' @NLP.label: 'Partner type code: 1=person, 2=organisation. NOT a name, never use for name lookups' BU_TYPE : String(2);</CODE></PRE><P>Without that label, <CODE>BU_TYPE</CODE> is just a two-character string column to the LLM, and "type code" columns get mistaken for name columns constantly, because the LLM has no way to know what the values mean. The label isn't decoration; it's the only channel you have to tell the model what NOT to do with a column. The same mechanism handles join direction (<CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1790377">@nlp</a>.joinType: 'LEFT'</CODE> when cardinality alone is ambiguous) and aliasing (<CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1790377">@nlp</a>.alias</CODE>).</P><P>None of this touches your OData service or Fiori UI. <CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1790377">@nlp</a>.label</CODE> is a CDS annotation the server reads at schema-load time and nothing else consumes.</P><HR /><H2 id="toc-hId-246168194">Configuration</H2><PRE><CODE>npm install @shahid.la/cds-db-nlquery-mcp</CODE></PRE><P>Create <CODE>.mcp.json</CODE> in your CAP project root:</P><PRE><CODE>{ "mcpServers": { "cds-db-nlquery-mcp": { "command": "npx", "args": ["-y", "@shahid.la/cds-db-nlquery-mcp"], "cwd": "/absolute/path/to/your/cap/project", "env": { "ANTHROPIC_API_KEY": "sk-ant-...", "ANTHROPIC_MODEL": "claude-haiku-4-5-20251001" } } } }</CODE></PRE><P>Open the project in an MCP-aware client and ask a question. <CODE>LLM_PROVIDER</CODE> auto-detects from whichever API key is set, so for a quick start you only need the key.</P><P><STRONG>Production env vars:</STRONG></P><P>Variable Default Purpose</P><TABLE><TBODY><TR><TD><CODE>MCP_ALLOWED_ENTITIES</CODE></TD><TD>all entities</TD><TD>Comma-separated entity allowlist, enforced on JOINs too</TD></TR><TR><TD><CODE>MCP_BLOCKED_COLUMNS</CODE></TD><TD>none</TD><TD>Columns stripped before the query is built</TD></TR><TR><TD><CODE>MCP_MAX_ROWS</CODE></TD><TD>500</TD><TD>Hard SQL <CODE>LIMIT</CODE> cap, server-side</TD></TR><TR><TD><CODE>MCP_DB_USER</CODE> / <CODE>MCP_DB_PASSWORD</CODE></TD><TD>inherits app's DB user</TD><TD>Run as a separate, ideally read-only, HANA user</TD></TR><TR><TD><CODE>MCP_MODEL_PATH</CODE></TD><TD><CODE>db</CODE></TD><TD>Where your CDS model lives, if not the default <CODE>db/</CODE></TD></TR></TBODY></TABLE><P><STRONG>LLM providers</STRONG>, today: Anthropic and OpenAI are first-class (<CODE>ANTHROPIC_API_KEY</CODE> / <CODE>OPENAI_API_KEY</CODE>). Everything else, Azure OpenAI, Groq, Ollama, local models, works through <CODE>OPENAI_BASE_URL</CODE> pointed at any OpenAI-compatible endpoint, not through a dedicated integration.</P><HR /><H2 id="toc-hId-49654689">What the Server Reads From Your CDS Schema</H2><P>Beyond <CODE>@title</CODE>/<CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1790377">@nlp</a>.label</CODE> (covered above), several more things already in your schema do the rest of the work, none of them new annotations you'd have to add just for this:</P><UL><LI><STRONG>Associations</STRONG>: drive every JOIN. Cardinality (<CODE>to-many</CODE> vs <CODE>to-one</CODE>) decides <CODE>LEFT</CODE> vs <CODE>INNER</CODE> automatically, and which <CODE>expand</CODE>/<CODE>hierarchy</CODE> shapes are even valid.</LI><LI><STRONG><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1648847">@Common</a>.Text</CODE> and <CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1648847">@Common</a>.ValueList</CODE></STRONG>: the standard SAP value-help patterns, reused so the LLM filters on human text instead of guessing codes, whether it's a small fixed enum or a large lookup table.</LI><LI><STRONG>Native CDS <CODE>enum</CODE></STRONG>: surfaced to the LLM as <CODE>name="value"</CODE> pairs, and translated back (<CODE>STATUS_text</CODE>) automatically in every result row, including inside nested <CODE>expand</CODE> results, recursively.</LI><LI><STRONG>Calculated-on-read elements</STRONG> (<CODE>FULL = FIRST || ' ' || LAST</CODE><span class="lia-unicode-emoji" title=":disappointed_face:">😞</span> selectable like any stored column. The server substitutes the underlying expression itself rather than assuming the database materialized it as a physical column (confirmed it doesn't, on HANA).</LI><LI><STRONG><CODE>@assert.range</CODE></STRONG>: surfaced as a sanity-bound hint, so the LLM can catch a likely unit/scale mismatch (a 0-1 ratio vs. a 0-100 percentage) before it emits a filter that trivially returns zero rows.</LI><LI><STRONG><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1434188">@CDS</a>.search</CODE></STRONG> and <STRONG><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1434188">@CDS</a>.valid.from</CODE>/<CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1434188">@CDS</a>.valid.to</CODE></STRONG>: free-text search columns and temporal entities, both read directly from annotations you'd already have for Fiori search bars and time-sliced data.</LI></UL><P>This is why it doesn't feel like a generic SQL wrapper bolted onto your project: it's reading the same metadata your CDS model already carries for OData, Fiori value-help, and UI labels, just pointed at a different consumer.</P><P>To make this concrete, here's the exact text <CODE>buildSchemaPrompt()</CODE> produces for the <A href="https://github.com/shahidla/cds-db-nlquery-mcp/tree/main/examples/capability-demo" target="_blank" rel="noopener nofollow noreferrer">capability demo schema</A>. This is what the LLM planner sees for every question:</P><PRE><CODE>Customers [Customers] columns: ID:String, NAME:String["Customer Name"], NOTES:String, FIRST:String, LAST:String, FULL:String[calculated] joins: "orders"→Orders(ID=CUSTOMER_ID,LEFT,toMany) searchable: NAME, NOTES Orders [Orders] columns: ID:String, CUSTOMER_ID:String, AMOUNT:Decimal{pairs with CURRENCY: always select both together}, CURRENCY:String["Currency code (e.g. USD, EUR). A text code, never numeric. Never SUM/AVG/MIN/MAX this column, aggregate AMOUNT instead."], STATUS:String{values: open="O",closed="C". Use the raw value in filters. Selecting "STATUS" alone ALSO gets you a "STATUS_text" business-term field in every result row, with NO extra effort: do not select "STATUS_text" yourself (it is not a real column, you cannot select it, it just appears in the output), and do not add your own "caseWhen" to relabel "STATUS" (you would create a duplicate/conflicting column with the one already added for you)}, ORDER_DATE:Date joins: "customer"→Customers(CUSTOMER_ID=ID,INNER), "items"→OrderItems(ID=ORDER_ID,LEFT,toMany) OrderItems [OrderItems] columns: ID:String, ORDER_ID:String, PRODUCT_ID:String, PRODUCT:String, QTY:Integer, STATUS:String{values: pending="P",shipped="S". Use the raw value in filters. Selecting "STATUS" alone ALSO gets you a "STATUS_text" business-term field in every result row, with NO extra effort: do not select "STATUS_text" yourself (it is not a real column, you cannot select it, it just appears in the output), and do not add your own "caseWhen" to relabel "STATUS" (you would create a duplicate/conflicting column with the one already added for you)} joins: "product"→Products(PRODUCT_ID=ID,INNER) Products [Products] columns: ID:String, NAME:String, SECRET:String, STATUS:String{values: active="A",discontinued="D". Use the raw value in filters. Selecting "STATUS" alone ALSO gets you a "STATUS_text" business-term field in every result row, with NO extra effort: do not select "STATUS_text" yourself (it is not a real column, you cannot select it, it just appears in the output), and do not add your own "caseWhen" to relabel "STATUS" (you would create a duplicate/conflicting column with the one already added for you)} Accounts [Accounts] columns: ID:String, NAME:String, PARENT_ID:String, STATUS:String{values: active="A",closed="X". Use the raw value in filters. Selecting "STATUS" alone ALSO gets you a "STATUS_text" business-term field in every result row, with NO extra effort: do not select "STATUS_text" yourself (it is not a real column, you cannot select it, it just appears in the output), and do not add your own "caseWhen" to relabel "STATUS" (you would create a duplicate/conflicting column with the one already added for you)} joins: "parent"→Accounts(PARENT_ID=ID,INNER){self-referencing: hierarchy}, "children"→Accounts(ID=PARENT_ID,LEFT,toMany){self-referencing: hierarchy} Sectors [Sectors] columns: CODE:String, DESCRIPTION:String Loans [Loans] columns: ID:String, DTI:Decimal[0..50], SECTOR:String{readable text available via "sector.DESCRIPTION". Include it in select to show the human-readable value, AND use this path (not the raw "SECTOR" column) when the question filters by a human term like "active"/"closed"/"overdue" rather than a raw code} joins: "sector"→Sectors(SECTOR=CODE,INNER) WorkAssignments [WorkAssignments] [temporal: valid from validFrom to validTo] columns: ID:String, EMPLOYEE:String, ROLE:String, validFrom:Date, validTo:Date</CODE></PRE><P>Entity names, column types, NLP labels, enum values with the auto-<CODE>STATUS_text</CODE> instruction, join cardinality, the <CODE>{self-referencing: hierarchy}</CODE> marker that enables hierarchy traversal, and the <CODE>[temporal: ...]</CODE> marker that enables <CODE>asOf</CODE> time-travel reads. The LLM doesn't infer any of this from column names. It reads it directly from what the schema reader extracted.</P><HR /><H2 id="toc-hId-200395541">Three Real Examples</H2><P>These run against the actual schema and seed data of <A href="https://github.com/shahidla/Banking-Sentinel" target="_blank" rel="noopener nofollow noreferrer">Banking Sentinel</A> (<A href="https://community.sap.com/t5/technology-blog-posts-by-sap/banking-sentinel-a-multi-agent-credit-risk-copilot-on-sap-btp-built-on-a/ba-p/14426538" target="_blank">writeup</A>), a demo SAP CAP + HANA Cloud project I use to exercise this package against a real, non-trivial schema (18 entities, multi-hop associations, coded value-help tables). Every descriptor and result below reflects that real schema. Nothing is invented for the post.</P><HR /><H3 id="toc-hId--289520971">Example 1: Simple Filter Across One Association</H3><P><STRONG>Question:</STRONG> <EM>"Which customers have a DTI ratio above 5?"</EM></P><P><STRONG>Schema (the relevant slice):</STRONG></P><PRE><CODE>entity BCA_DTI { key PARTNER : String(10); DTI_RATIO : Decimal(5,2); customer : Association to BusinessPartners on customer.PARTNER = PARTNER; } @NLP.label: 'Customers and business partners. Demo borrowers: 301xxxx, guarantors: 309xxxx' entity BusinessPartners { key PARTNER : String(10); @title: 'Customer / Business Partner Name' BU_SORT1 : String(50); }</CODE></PRE><P><STRONG>Descriptor the LLM returns:</STRONG></P><PRE><CODE>{ "entity": "BCA_DTI", "select": ["PARTNER", "DTI_RATIO", "customer.BU_SORT1"], "where": [{ "col": "DTI_RATIO", "op": "&gt;", "val": 5 }], "orderBy": "DTI_RATIO", "orderDir": "DESC", "limit": 20 }</CODE></PRE><P><STRONG>Result, against the live database:</STRONG></P><PRE><CODE>Results: 3 rows Partner : 30100003 DTI Ratio : 7.20 Customer : Domestic Customer AU 3 Partner : 30100001 DTI Ratio : 5.80 Customer : Domestic Customer AU 1 Partner : 30100004 DTI Ratio : 5.40 Customer : Domestic Customer AU 4</CODE></PRE><P>One association (<CODE>customer</CODE>), one JOIN, one filter. This is the floor. Every query, however complex, starts from this same mechanism.</P><HR /><H3 id="toc-hId--486034476">Example 2: Coded Values via <CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1648847">@Common</a>.Text</CODE> (Not Guessing the Raw Code)</H3><P><STRONG>Question:</STRONG> <EM>"Show me active loans for customers in the mining sector, with the borrower's name and loan amount."</EM></P><P>This question has a trap: "active" and "mining sector" are both <EM>human</EM> terms over <EM>coded</EM> columns. <CODE>Loans.STATUS</CODE> is stored as a single character (<CODE>'A'</CODE>/<CODE>'C'</CODE>), not the word "active." If the LLM guessed at the raw code, it would be wrong as often as it was right.</P><P><STRONG>Schema:</STRONG></P><PRE><CODE>entity Loans { key LOAN_ID : String(15); PARTNER : String(10); AMOUNT : Decimal(15,2); @Common.Text: status.TEXT STATUS : String(1); customer : Association to BusinessPartners on customer.PARTNER = PARTNER; status : Association to LoanStatusCodes on status.CODE = STATUS; } // Adding a new status is a data INSERT into this table, never a schema/code change. entity LoanStatusCodes { key CODE : String(1); TEXT : String(20); } entity BusinessPartners { key PARTNER : String(10); BU_SORT1 : String(50); SECTOR_CODE : String(20); // RETAIL_PROP, MINING, AGRICULTURE, ... }</CODE></PRE><P><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1648847">@Common</a>.Text: status.TEXT</CODE> is the standard SAP value-help pattern, the same one Fiori uses to show "Active" in a dropdown while storing <CODE>'A'</CODE>. The schema reader tells the LLM this path exists; the system prompt instructs it to filter on <CODE>status.TEXT</CODE> directly rather than invent a raw code.</P><P><STRONG>Descriptor the LLM returns:</STRONG></P><PRE><CODE>{ "entity": "Loans", "select": ["LOAN_ID", "AMOUNT", "customer.BU_SORT1", "status.TEXT"], "where": [ { "col": "customer.SECTOR_CODE", "op": "=", "val": "MINING" }, { "col": "status.TEXT", "op": "like", "val": "Active" } ], "limit": 50 }</CODE></PRE><P><STRONG>Result:</STRONG></P><PRE><CODE>Results: 1 row Loan ID : L-009 Amount : AUD 3,200,000.00 Customer : Domestic Customer AU 9 Status : Active</CODE></PRE><P>If the bank adds a new loan status next quarter (<CODE>PENDING_REVIEW = 'P'</CODE>), this query keeps working without a code change. It's matching on text, not a hardcoded code the LLM would otherwise have had to memorize.</P><HR /><H3 id="toc-hId--682547981">Example 3: Column-to-Column Comparison Across Two Hops (<CODE>valCol</CODE>)</H3><P><STRONG>Question:</STRONG> <EM>"Which loans are under-collateralized, where the pledged collateral is worth less than the loan amount?"</EM></P><P>This is qualitatively different from the first two: it's not comparing a column to a value the user typed, it's comparing <STRONG>two columns from different entities</STRONG> to each other. There's no literal to filter on.</P><P><STRONG>Schema:</STRONG></P><PRE><CODE>entity BCA_COLLATERAL { key LOAN_ID : String(15); key COLLAT_ID : String(15); COLLAT_TYPE : String(10); // PROPERTY, VEHICLE, CASH VALUE : Decimal(15,2); loan : Association to Loans on loan.LOAN_ID = LOAN_ID; } entity Loans { key LOAN_ID : String(15); AMOUNT : Decimal(15,2); customer : Association to BusinessPartners on customer.PARTNER = PARTNER; }</CODE></PRE><P>The descriptor format supports <CODE>valCol</CODE> for exactly this: comparing a column to another column (a path, possibly through a JOIN) instead of to a literal <CODE>val</CODE>:</P><PRE><CODE>{ "entity": "BCA_COLLATERAL", "select": ["COLLAT_TYPE", "VALUE", "loan.LOAN_ID", "loan.PARTNER", "loan.customer.BU_SORT1", "loan.AMOUNT"], "where": [ { "col": "VALUE", "op": "&lt;", "valCol": "loan.AMOUNT" } ], "orderBy": "loan.LOAN_ID", "limit": 50 }</CODE></PRE><P><CODE>loan.customer.BU_SORT1</CODE> is a two-hop path (<CODE>BCA_COLLATERAL → loan → Loans → customer → BusinessPartners</CODE>), resolved in the same single query, no extra round-trip. The generated <CODE>WHERE</CODE> clause is <CODE>WHERE COLLAT.VALUE &lt; L.AMOUNT</CODE>, a real cross-table comparison, not two separate queries reconciled in JavaScript.</P><P><STRONG>Raw result: 10 rows.</STRONG> Worth pausing on why it's 10 and not fewer: the comparison is <STRONG>per pledged asset</STRONG>, not per loan. A loan secured by both a property and a cash deposit produces two collateral rows, and each is compared to the <EM>full</EM> loan amount individually. That's the literal, correct answer to "which pledged assets are worth less than the loan." It just means a multi-asset loan can appear more than once, even when its combined collateral would be adequate.</P><P>That's exactly where Stage 3 earns its place. The MCP server hands back those 10 flat rows; the second LLM call (the one that turns MCP results into an answer) grouped them by loan and computed the shortfall itself:</P><PRE><CODE>10 under-collateralized loan records identified (collateral value &lt; loan amount): Loan L-004 (Partner: 30100003, Domestic Customer AU 3) Loan Amount: AUD 2,100,000.00 Total Collateral Value: AUD 1,480,000.00 Shortfall: AUD 620,000.00 Collateral Breakdown: Property: AUD 1,200,000.00 Cash: AUD 280,000.00 Loan L-006 (Partner: 30100005, Domestic Customer AU 5) Loan Amount: AUD 1,850,000.00 Total Collateral Value: AUD 1,300,000.00 Shortfall: AUD 550,000.00 Collateral Breakdown: Property: AUD 1,100,000.00 Cash: AUD 200,000.00 Loan L-007 (Partner: 30100006, Domestic Customer AU 6) Loan Amount: AUD 45,000.00 Total Collateral Value: AUD 35,000.00 Shortfall: AUD 10,000.00 Collateral Breakdown: Vehicle: AUD 35,000.00 Loan L-009 (Partner: 30100009, Domestic Customer AU 9) Loan Amount: AUD 3,200,000.00 Total Collateral Value: AUD 2,400,000.00 Shortfall: AUD 800,000.00 Collateral Breakdown: Property: AUD 2,000,000.00 Cash: AUD 400,000.00 [+ 3 more collateral rows across 2 loans from the wider training portfolio, outside the named demo customers]</CODE></PRE><P>This example is the most demanding of the three: it resolves a coded value-help table and performs a column-to-column comparison across an association path, which requires the framework, not the LLM, to understand JOIN semantics. And the per-asset-not-per-loan result is itself a useful, honest reminder of what the descriptor format can and can't express today: it has no <CODE>SUM</CODE>/<CODE>GROUP BY</CODE>, so "is this loan's <EM>combined</EM> collateral sufficient" is a question for the LLM answering over the raw rows, not something the query itself computes.</P><HR /><H2 id="toc-hId--585658479">Beyond the Three Examples: Four More Capabilities With Real Output</H2><P>The Banking Sentinel examples cover three patterns: simple filter, coded value-help JOIN, and column-to-column comparison. Here's what else the descriptor format supports, each shown against the <A href="https://github.com/shahidla/cds-db-nlquery-mcp/tree/main/examples/capability-demo" target="_blank" rel="noopener nofollow noreferrer">capability demo schema</A> with output captured by running <CODE>node examples/capability-demo/generate.js</CODE> against a real in-memory database.</P><HR /><H3 id="toc-hId--1075574991">Nested Reads</H3><P><STRONG>Question:</STRONG> <EM>"Orders with their line items nested inside"</EM></P><PRE><CODE>{ "entity": "Orders", "select": ["ID", "CUSTOMER_ID"], "expand": [{ "assoc": "items", "select": ["PRODUCT", "QTY"] }] }</CODE></PRE><P><STRONG>Result:</STRONG></P><PRE><CODE>[ { "ID": "O1", "CUSTOMER_ID": "C1", "items": [{"PRODUCT":"Widget","QTY":10},{"PRODUCT":"Gadget","QTY":5}] }, { "ID": "O2", "CUSTOMER_ID": "C1", "items": [{"PRODUCT":"Widget","QTY":20}] }, { "ID": "O3", "CUSTOMER_ID": "C2", "items": [{"PRODUCT":"Gizmo","QTY":3}] }, { "ID": "O4", "CUSTOMER_ID": "C2", "items": [] }, ... ]</CODE></PRE><P>One query. Each parent row carries its children as a real nested array, not a flattened, duplicated-parent-row JOIN result. <CODE>expand</CODE> supports <CODE>orderBy</CODE>/<CODE>limit</CODE> on the nested side (e.g. "each order's single largest line item by quantity") and nests recursively to any depth.</P><HR /><H3 id="toc-hId--1272088496">Recursive Hierarchy</H3><P><STRONG>Question:</STRONG> <EM>"All descendants of account A1, the full org tree below it"</EM></P><PRE><CODE>{ "entity": "Accounts", "select": ["ID", "NAME", "PARENT_ID"], "hierarchy": { "assoc": "children", "direction": "descendants", "startWhere": [{ "col": "ID", "op": "=", "val": "A1" }] } }</CODE></PRE><P><STRONG>Result:</STRONG></P><PRE><CODE>ID : A1 Name : Holding Co Parent : — Status : active ID : A2 Name : Regional Division Parent : A1 Status : active ID : A3 Name : Local Branch North Parent : A2 Status : active ID : A4 Name : Local Branch South Parent : A2 Status : closed ID : A5 Name : Sub Branch North-1 Parent : A3 Status : active</CODE></PRE><P>Five levels from a single question. The <CODE>STATUS_text</CODE> translation (raw <CODE>"A"</CODE> to <CODE>"active"</CODE>, <CODE>"X"</CODE> to <CODE>"closed"</CODE>) applies automatically, including on hierarchy results. A fixed-depth association path (<CODE>account.parent.parent.NAME</CODE>) cannot express an unbounded tree walk. <CODE>hierarchy</CODE> can, capped by a configurable max depth.</P><HR /><H3 id="toc-hId--1468602001">Window Functions</H3><P><STRONG>Question:</STRONG> <EM>"Each customer's single largest order"</EM></P><PRE><CODE>{ "entity": "Orders", "select": ["ID", "CUSTOMER_ID", "AMOUNT"], "window": [{ "fn": "rank", "as": "RANK", "partitionBy": ["CUSTOMER_ID"], "orderBy": [{ "col": "AMOUNT", "dir": "DESC" }] }], "windowFilter": [{ "col": "RANK", "op": "=", "val": 1 }] }</CODE></PRE><P>CDS generates the subquery wrapping that <CODE>HAVING</CODE> alone can't express:</P><PRE><CODE>SELECT ID, CUSTOMER_ID, AMOUNT, RANK FROM ( SELECT ID, CUSTOMER_ID, AMOUNT, rank() OVER (PARTITION BY CUSTOMER_ID ORDER BY AMOUNT DESC) AS RANK FROM Orders ) WHERE RANK = 1 LIMIT 50</CODE></PRE><P><STRONG>Result:</STRONG></P><PRE><CODE>ID : O2 Customer : C1 (Acme Corp) Amount : 2300.50 ID : O4 Customer : C2 (Globex Inc) Amount : 4200.00 ID : O5 Customer : C3 (Initech) Amount : 150.00</CODE></PRE><P>One row per customer, the actual largest order, not a <CODE>MAX()</CODE> that loses the order ID. Running totals, lag/lead, and <CODE>row_number</CODE> follow the same pattern.</P><HR /><H3 id="toc-hId--1665115506">Time-Travel Reads</H3><P>Two questions, same entity, different answers:</P><P><STRONG>"What was Alice's role on 2026-02-15?"</STRONG></P><PRE><CODE>{ "entity": "WorkAssignments", "select": ["EMPLOYEE", "ROLE"], "where": [{ "col": "EMPLOYEE", "op": "=", "val": "Alice" }], "asOf": "2026-02-15" }</CODE></PRE><P>→ <CODE>ROLE: Analyst</CODE></P><P><STRONG>"What is Alice's current role?"</STRONG></P><PRE><CODE>{ "entity": "WorkAssignments", "select": ["EMPLOYEE", "ROLE"], "where": [{ "col": "EMPLOYEE", "op": "=", "val": "Alice" }] }</CODE></PRE><P>→ <CODE>ROLE: Senior Analyst</CODE></P><P><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1434188">@CDS</a>.valid.from</CODE>/<CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1434188">@CDS</a>.valid.to</CODE> entities work automatically. The server adds the <CODE>validFrom &lt;= date AND validTo &gt; date</CODE> filter for <CODE>asOf</CODE>, or defaults to now when no date is given. No annotation beyond what you'd already add for Fiori time-sliced data.</P><HR /><P>Same architectural promise as before: the LLM picks which of these to use and fills in the JSON, CDS turns it into real CQN, the server never writes SQL by hand.</P><HR /><H2 id="toc-hId--1568226004">Built by Testing Against Real HANA, Not Just an In-Memory Database</H2><P>A comprehensive unit test suite (126 tests) passed the entire time these capabilities were broken in a specific, real way. The tests used an in-memory SQLite-style adapter that tolerates things real HANA doesn't. The only way to actually know whether this worked was to deploy the demo schema to a real HANA Cloud instance and run every query against it for real. So that's what I did, repeatedly, across one extended session, and it caught real bugs:</P><UL><LI>A real NL question, "show me each customer's single largest order," surfaced that <CODE>expand</CODE>'s <CODE>orderBy</CODE>/<CODE>limit</CODE> (for "top N per group") wasn't implemented at all. The row cap applied, but nothing sorted first, so the result could silently be an arbitrary order instead of the actual largest one. Fixed by sorting before truncating. That fix, plus the existing enum-to-text translation and blocked-column stripping, all shared one unexamined assumption: each only ever checked <CODE>Array.isArray()</CODE> on a nested value. A follow-up systematic audit, deliberately constructing a two-level <CODE>expand</CODE> test, not a natural-language question this time, found all three silently skipped a nested value entirely whenever it was a plain object (a <CODE>to-one</CODE> association) instead of an array.</LI><LI>A descriptor with an unrecognized field (an LLM wrote <CODE>{"notExists": "orders"}</CODE> as a sibling of <CODE>"where"</CODE> instead of inside it) was silently ignored rather than rejected. The query ran with no filter applied at all, returning every row instead of failing loudly.</LI><LI>The legacy <CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/2302137">@SAP</a>/hana-client</CODE>-based HANA runtime (still the default most CAP+HANA projects use) silently drops the <CODE>OVER</CODE> clause from a window-function query. No error, just wrong SQL, until HANA itself throws a syntax error downstream. Confirmed the CQN this package generates is <EM>correct</EM> per CAP's current standard (the same shape <CODE>@cap-js/sqlite</CODE> and the modern <CODE>@cap-js/hana</CODE> adapter both render properly). Only the legacy runtime mishandles it. The server now detects which adapter is connected and only restricts behavior on the legacy one; on a modern adapter, the same query just works.</LI><LI>A genuinely production question, Banking Sentinel's own "what is the total loan amount across all customers?", surfaced that the planning LLM occasionally puts a function-call string like <CODE>"SUM(AMOUNT)"</CODE> directly into <CODE>select</CODE>, instead of using the <CODE>aggregate</CODE> field. The column-ref builder then treated the whole string as a literal column name, and HANA rejected it with a cryptic <CODE>invalid column name: SUM(AMOUNT)</CODE>. A second real failure on the same question widened the same mistake to a full SQL fragment with a trailing alias (<CODE>"SUM(AMOUNT) AS TOTAL_AMOUNT"</CODE>), which an exact-match pattern didn't catch. Both are now rejected up front, across every field a column-spec string can appear in (<CODE>select</CODE>, <CODE>where</CODE>, <CODE>groupBy</CODE>, <CODE>orderBy</CODE>, <CODE>having.col</CODE>, <CODE>aggregate.col</CODE>), with a message that names the actual fix, instead of reaching HANA as a cryptic failure or, worse, silently returning a wrong total.</LI></UL><P>Every one of these was found, root-caused, and fixed by actually deploying <CODE>examples/capability-demo/</CODE>'s schema to live HANA and running real queries against it, not by reading the code harder. The four scripts that do this live in that folder in the GitHub repo (they're development/testing tools, not part of the published npm package, <CODE>npm install</CODE> only gives you <CODE>src/</CODE>), ready to run, not just described:</P><UL><LI><CODE>validate-deployment.js</CODE>: runs every hand-written descriptor against your deployment and compares rows to the SQLite-generated <CODE>results.json</CODE> baseline, no LLM involved. The definitive check for whether execution is correct on your specific backend and adapter.</LI><LI><CODE>ask.js</CODE>: runs a single natural-language question end-to-end through a real LLM against your deployment, useful for questions the pre-built list doesn't cover.</LI><LI><CODE>ask-batch.js</CODE>: runs every pre-built question's natural-language text (not the hand-written descriptor) through a real LLM and compares the result. The script that actually distinguishes "the package is wrong" from "the model picked an odd column," and the one that surfaced real, fixable mistakes across two different models (Claude and DeepSeek).</LI><LI><CODE>smoke-test-server.js</CODE>: unlike the other three (which call internal functions via <CODE>require()</CODE>), this spawns <CODE>src/mcp-server.js</CODE> as a real child process and drives it through the MCP stdio protocol the same way Claude Code or <CODE>npx</CODE> actually would. Used to confirm that what's correct at the library level is also correct at the published-artifact level.</LI></UL><P>If you're evaluating this for something that matters, don't take the examples above on faith. Clone the repo, deploy <CODE>examples/capability-demo/</CODE> to your own HANA, and run these scripts yourself.</P><HR /><H2 id="toc-hId--1596555818">What Happens When the LLM Gets It Wrong</H2><P>All three examples above are success cases, worth being honest about the failure path too. If the LLM's response doesn't parse as JSON, or comes back without an <CODE>entity</CODE> field, the server rejects it outright rather than guessing or silently running a degraded query. The same applies to an entity or column outside <CODE>MCP_ALLOWED_ENTITIES</CODE>, or one that doesn't exist in your schema at all. Each produces a clear error back to the MCP client, not a best-effort result you'd have to double-check. There's no fallback path that quietly does something different from what you asked.</P><HR /><H2 id="toc-hId--1793069323">When to Use This (and When Not To)</H2><P><STRONG>Good fit:</STRONG></P><UL><LI>Ad-hoc data exploration during development or support investigations</LI><LI>Audit/compliance spot-checks, e.g. "show me all loans without collateral where LTV exceeds 80%"</LI><LI>Operational questions that change shape every time, so a fixed report doesn't fit</LI></UL><P><STRONG>Not a fit:</STRONG></P><UL><LI>Anything user-facing: this bypasses <CODE>@requires</CODE>/<CODE>@restrict</CODE>, use your OData service for that</LI><LI>Scheduled/repeated reporting: this is exploratory, not a BI tool</LI><LI>Any write path: read-only by design, no exceptions</LI></UL><P><STRONG>Database compatibility:</STRONG></P><UL><LI>Tested on HANA Cloud (BTP) with both adapters CDS supports: <CODE>@cap-js/hana</CODE> (all capabilities work, including window functions and <CODE>viaFiltered</CODE>-inside-aggregate) and the legacy <CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/2302137">@SAP</a>/hana-client</CODE> runtime (those two specific things are rejected with a clear error; everything else works). Not tested on on-premise classic HANA.</LI></UL><HR /><H2 id="toc-hId--1989582828">Try It</H2><PRE><CODE>npm install @shahid.la/cds-db-nlquery-mcp</CODE></PRE><P>Add the <CODE>.mcp.json</CODE> block above, point it at a CAP project with a few <CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1790377">@nlp</a>.label</CODE> annotations, and ask it something you'd otherwise have opened a SQL editor for.</P><P>Want to verify all of this yourself before trusting it with something that matters, rather than trusting this post? Clone the repo, deploy <CODE>examples/capability-demo/</CODE>'s schema to your own database, and run <CODE>validate-deployment.js</CODE>, <CODE>ask.js</CODE>, <CODE>ask-batch.js</CODE>, and <CODE>smoke-test-server.js</CODE>, the same scripts that found and confirmed every fix described above.</P><P data-unlink="true">Find the project on <STRONG><A href="https://github.com/shahidla/cds-db-nlquery-mcp" target="_blank" rel="noopener nofollow noreferrer">GitHub</A></STRONG>, <STRONG><A href="https://www.npmjs.com/package/@shahid.la/cds-db-nlquery-mcp" target="_blank" rel="noopener nofollow noreferrer">npm</A></STRONG>, and the <STRONG><A href="https://registry.modelcontextprotocol.io/?q=cds-db-nlquery-mcp" target="_self" rel="nofollow noopener noreferrer">MCP Registry</A>&nbsp;</STRONG>.</P><P>Feedback, issues, and PRs welcome, especially real-world schema patterns and annotation edge cases.</P> 2026-06-25T03:37:23.340000+02:00 https://community.sap.com/t5/artificial-intelligence-blogs-posts/bringing-domain-intelligence-into-sap-joule/ba-p/14431704 Bringing Domain Intelligence into SAP Joule 2026-07-01T23:25:41.416000+02:00 cassiobinkowski https://community.sap.com/t5/user/viewprofilepage/user-id/4955 <H2 id="toc-hId-1819049836">Why “domain intelligence” is the missing piece</H2><P>Foundation models are impressive generalists, but enterprise questions are specific. “Which quotation gives us the best win/win with this customer?” is not answered by general world knowledge. It is answered by <EM>your</EM> master data, <EM>your&nbsp;</EM>pricing logic, and the relationships between them.</P><P>That is where two capabilities on SAP Business AI Platform come together. First, a <STRONG>Knowledge Graph</STRONG> captures your domain as connected, machine-readable facts with explicit semantics. Second, the <STRONG>A2A protocol</STRONG> lets a custom agent that reasons over that graph plug directly into Joule, so the intelligence shows up where users already work.</P><P>The picture below shows how the pieces fit: Joule stays the orchestrator, a pro-code agent does the specialized work, and the SAP HANA Cloud Knowledge Graph engine holds the domain intelligence the agent draws on.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Architecture - Custom Intelligence in Joule.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428371iDEB49CDD5264C438/image-size/large?v=v2&amp;px=999" role="button" title="Architecture - Custom Intelligence in Joule.png" alt="Architecture - Custom Intelligence in Joule.png" /></span></P><P><EM>Joule orchestrates; the pro-code agent reasons; the Knowledge Graph engine holds your domain intelligence.</EM></P><H2 id="toc-hId-1622536331">Step 1: Capture domain knowledge as an Ontology and Knowledge Graph</H2><P>A Knowledge Graph is only as smart as the <STRONG>ontology</STRONG> behind it. The ontology is the semantic model: it names your entities, describes their attributes, and, crucially, records how they relate. That semantic layer is what turns raw content into meaning a machine can reason over.</P><P>Importantly, your source does not have to be a database. The SAP HANA Cloud Knowledge Graph samples cover two paths. In the first, <STRONG>unstructured documents</STRONG> such as PDFs are turned into a graph: a large language model extracts the entities and the relationships between them from the text and converts them into RDF triples. In the second, <STRONG>tabular data</STRONG> is modeled as an ontology from its schema, keys and joins. Either way, the result is serialized to <STRONG>RDF in Turtle (.ttl)</STRONG>, so the same downstream steps apply regardless of where the knowledge came from.</P><P>Loading it into SAP HANA Cloud is straightforward. You ingest the triples through the SPARQL_EXECUTE procedure, or stage the file in cloud storage and pull it in with IMPORT FROM RDF FILE from the Database Explorer. Once ingested, the graph is queryable with <STRONG>SPARQL</STRONG>, including from natural language, so business users can ask a question in plain words and get a grounded, relationship-aware answer.</P><P>The pipeline looks like this, from your source data on the left to a query-ready graph on the right. The labels are deliberately generic; your entities, documents and agent will be your own:</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Flow - From Data to Dialogue.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428372i2F8EE327B92B51F2/image-size/large?v=v2&amp;px=999" role="button" title="Flow - From Data to Dialogue.png" alt="Flow - From Data to Dialogue.png" /></span></P><P><EM>Build the graph once, then reuse it across pro-code agents and every Joule conversation.</EM></P><P>The important idea for partners: you model the graph once. The same semantic asset can then ground many agents and many conversations, rather than being rebuilt for each use case.</P><H2 id="toc-hId-1426022826">Step 2: Bring your agent into Joule with A2A</H2><P>Joule acts as the central orchestrator and entry point. When a user makes a request, Joule’s planning and reasoning engine decides how to fulfil it, which can include delegating the task to a custom agent. SAP supports two integration patterns, and the right one depends on how the agent was built.</P><P><STRONG>For low-code agents built with Joule Studio</STRONG>, integration is largely automatic. Deploying the agent generates the required Joule artifacts, including a Joule Scenario and a Dialog Function, and registers the scenario in Joule’s Scenario Catalog. There is no endpoint or protocol to manage by hand.</P><P><STRONG>For pro-code agents</STRONG>, the “Bring Your Own Agent” (BYOA) pattern uses the open <STRONG>A2A protocol</STRONG>, so you can build with any framework that speaks A2A. The flow is clean and decoupled:</P><OL><LI>Expose your agent as an <STRONG>A2A server</STRONG> with an HTTP endpoint that follows the A2A specification.</LI><LI>Create a <STRONG>Joule Scenario</STRONG> for the agent, and add a <STRONG>Dialog Function</STRONG> with an action of type agent-request.</LI><LI>Point that Dialog Function at your agent’s A2A endpoint.</LI></OL><P>At runtime, Joule becomes the <STRONG>A2A client</STRONG>. It prepares a request using the message/send method (A2A version 0.3.0), sends the user’s utterance to your agent, waits for the response, and presents the result. Joule handles synchronous replies within a 60 second window; for longer work, the agent uses <STRONG>push notifications</STRONG> to a Joule webhook, and <STRONG>context and task IDs</STRONG> carry state across multi-turn conversations. Secure communication is established through an <STRONG>Identity Authentication Service (IAS) App2App</STRONG> trust relationship between Joule and your agent.</P><P>One practical note worth setting expectations on: the Agent Gateway is not yet generally available, so today the architecture supports unidirectional, outbound communication from Joule to the agent. The bidirectional building blocks are expected to follow.</P><H2 id="toc-hId-1229509321">Two building blocks today, one coherent example soon</H2><P>It is worth being clear about where things stand. Today these are <STRONG>two separate reference assets</STRONG>, not one finished scenario. The Knowledge Graph samples show how to build and query a graph from your documents or tables. The pro-code agent reference, btp-joule-a2a-pro-code-agent, shows the other half: a production-grade agent that plugs into Joule over A2A.</P><P>That reference agent is written in <STRONG>TypeScript on CAP</STRONG>, uses <STRONG>LangGraph</STRONG> for its logic, reaches foundation models through the <STRONG>SAP Cloud SDK for AI</STRONG> and <STRONG>Generative AI Hub</STRONG>, and demonstrates the full A2A round trip with Joule, including human-in-the-loop and webhook-based asynchronous updates, with Terraform for the BAIP setup. A companion community walkthrough by Felix Bartler shows the minimal version of the same A2A wiring.</P><P>The target pattern joins the two halves: a pro-code agent that treats your <STRONG>HANA Cloud Knowledge Graph as a tool</STRONG>, issuing SPARQL queries to fetch precise, relationship-aware facts, then reasoning with Generative AI Hub before returning a grounded answer through Joule. The graph supplies the trustworthy facts; the model supplies the reasoning; A2A supplies the delivery.</P><H2 id="toc-hId-1032995816">What this unlocks for SAP partners</H2><P>This is where the partner opportunity is clearest. The pattern lets you help <STRONG>your customers bring their own domain data, knowledge and ontology into SAP</STRONG>, where it becomes a governed, reusable Knowledge Graph that grounds Joule. Instead of standing up a separate AI stack, your customers <STRONG>leverage the SAP platform investments they already have</STRONG>, including SAP HANA Cloud, SAP Business AI Platform (BAIP), Generative AI Hub and Joule, while you supply the domain expertise and the agent on top. Their proprietary knowledge stays in their SAP landscape, under their governance.</P><P>The advantages compound. The Knowledge Graph is a reusable asset that encodes domain expertise once and grounds every agent that uses it, which reduces hallucination and makes answers defensible. A2A is an open standard, so you keep framework freedom and avoid lock-in while still landing natively inside the SAP experience. And because the intelligence surfaces through Joule, adoption is frictionless: customers’ users get specialized answers in the assistant they already trust.</P><P>The build path is incremental. Start by turning one high-value slice of a customer’s knowledge, a set of documents or a key data model, into an ontology and Knowledge Graph in SAP HANA Cloud. Wrap a pro-code agent around it that queries the graph and reasons with Generative AI Hub. Then expose the agent over A2A and register it as a Joule Scenario. Each step is independently useful, and together they turn generic AI into <EM>your customer’s</EM> AI.</P><H2 id="toc-hId-836482311">Further resources to get you started</H2><UL><LI><STRONG><SPAN>Building Intelligent Data Applications with SAP HANA Cloud Knowledge Graphs:</SPAN></STRONG>&nbsp;<A href="https://discovery-center.cloud.sap/protected/index.html#/missiondetail/4568/4856/" target="_blank" rel="nofollow noopener noreferrer">https://discovery-center.cloud.sap/protected/index.html#/missiondetail/4568/4856/</A></LI><LI><STRONG>Knowledge Graph samples (documents and tables to ontology, 7 scenarios):<SPAN>&nbsp; </SPAN></STRONG><A href="https://github.com/IDGCOENA/saphanacloudkge" target="_blank" rel="noopener nofollow noreferrer">github.com/IDGCOENA/saphanacloudkge</A></LI><LI><STRONG>A2A pro-code agent reference implementation:<SPAN>&nbsp; </SPAN></STRONG><A href="https://github.com/SAP-samples/btp-joule-a2a-pro-code-agent" target="_blank" rel="noopener nofollow noreferrer">github.com/SAP-samples/btp-joule-a2a-pro-code-agent</A></LI><LI><STRONG>SAP Architecture Center, Integrating AI Agents with Joule:<SPAN>&nbsp; </SPAN></STRONG><A href="https://architecture.learning.sap.com/docs/ref-arch/ca1d2a3e/4" target="_blank" rel="noopener noreferrer">architecture.learning.sap.com</A></LI><LI><STRONG>Community walkthrough, Connect Code Based Agents into Joule:<SPAN>&nbsp; </SPAN></STRONG><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/joule-a2a-connect-code-based-agents-into-joule/ba-p/14329279" target="_blank">community.sap.com</A></LI><LI><STRONG>A2A protocol specification:<SPAN>&nbsp; </SPAN></STRONG><A href="https://a2a-protocol.org/latest/" target="_blank" rel="noopener nofollow noreferrer">a2a-protocol.org</A></LI></UL><H2 id="toc-hId-639968806">&nbsp;</H2><H2 id="toc-hId-443455301">Coming next - Stay tuned!</H2><P><EM><STRONG>Coming next.<SPAN>&nbsp; </SPAN></STRONG>We are building a coherent, end-to-end example that connects a custom Knowledge Graph to a Joule agent over A2A, and we are preparing a <STRONG>hands-on workshop</STRONG> to walk partners through it.</EM></P><P>Share your use cases and questions in the comments. Don't hesitate to reach out as well in case you want some more close support.</P> 2026-07-01T23:25:41.416000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/stop-tuning-your-database-manually-let-sap-hana-cloud-s-intelligent/ba-p/14422705 Stop Tuning Your Database Manually - Let SAP HANA Cloud's Intelligent Advisors Do It 2026-07-02T13:05:27.997000+02:00 jinheejeong https://community.sap.com/t5/user/viewprofilepage/user-id/159990 <DIV class=""><SPAN>Managing a production SAP HANA Cloud database means constantly juggling performance, cost, and complexity. Tables grow beyond limits. Indexes drift out of sync with query patterns. Workload peaks spike memory and compute. Knowing </SPAN><EM>what</EM><SPAN> to change - and </SPAN><EM>when</EM><SPAN>&nbsp;- has traditionally required deep expertise and hours of manual analysis.</SPAN></DIV><DIV class=""><P class="">That's changing. SAP HANA Cloud&nbsp;delivers a suite of built-in functional advisors that analyze your instance and surface specific, actionable recommendations - directly from SAP HANA Cloud Central.</P><!-- What are the advisors --><H2 id="toc-hId-1818156107"><FONT color="#000000">What Are the Functional Advisors?</FONT></H2><P>The advisors are specialized analysis engines embedded in SAP HANA Cloud Central, accessible through the <STRONG>Recommendations</STRONG> card on any instance's Overview page. Each advisor monitors a specific domain - physical design, storage tiering, compute scaling, instance sizing - and produces recommendations with ready-to-apply actions.</P><DIV class=""><DIV class=""><H3 id="toc-hId-1750725321"><FONT color="#3366FF">Partition Advisor</FONT></H3></DIV><DIV class=""><H4 id="toc-hId-1683294535">The problem it solves</H4><P>SAP HANA Cloud has a hard limit of 2 billion records per table. Large, unpartitioned tables also cause slow scans, data skew, and expensive delta merges. Knowing how to partition - which method, which key, how many partitions - requires deep knowledge of your data's shape and access patterns.</P><H4 id="toc-hId-1486781030">What the advisor does</H4><P>It examines table specifications, row counts, query access patterns, and column cardinality, then applies a rule set that covers full-table limits, full-partition limits, frequent query predicates, temporal access patterns, and dynamic partition overflow. For each table it produces a single consolidated SQL recommendation with an importance level (High, Medium, or Low) so you know where to act first.</P><H4 id="toc-hId-1290267525">Key benefits</H4><UL><LI>Proactively flags tables approaching the 2 billion record limit before they cause operational incidents</LI><LI>Recommends the right partitioning method (HASH, RANGE, Round-Robin, multi-level) based on actual data - not assumptions</LI><LI>Reduces the cost and risk of repartitioning large tables by getting the design right earlier</LI><LI>Works particularly well for very large tables, slow scan scenarios, data skew problems, and growing data volumes</LI></UL><DIV class=""><H4 id="toc-hId-1093754020">Getting started</H4>Enable the Partition Advisor from the Advisors tab. Provide credentials for a dedicated technical database user (not DBADMIN) with <CODE>CATALOG READ</CODE> and <CODE>PARTITION ADMIN</CODE> privileges. Default thresholds apply out of the box, but you can tune <CODE>MIN_ROWS_FOR_PARTITIONING</CODE>, <CODE>REPARTITIONING_THRESHOLD</CODE>, and <CODE>INITIAL_PARTITIONS</CODE> to match your operational standards.</DIV></DIV></DIV><!-- ── Index Advisor ── --><DIV class=""><DIV class=""><H3 id="toc-hId-768157796"><FONT color="#3366FF">Index Advisor</FONT></H3></DIV><DIV class=""><H4 id="toc-hId-700727010">The problem it solves</H4><P>Indexes that matched your query patterns six months ago may be dead weight today. Missing indexes leave performance on the table. Building the right index coverage manually means correlating workload statistics with table structures - a time-consuming process that rarely happens on any schedule.</P><H4 id="toc-hId-504213505">What the advisor does</H4><P>It analyzes all tables, existing index definitions, and workload statistics to identify two types of opportunities:</P><UL><LI><STRONG>Create index</STRONG> — where adding an index on a specific column is expected to improve query performance based on observed access patterns</LI><LI><STRONG>Drop index</STRONG> — where an index is rarely used or provides limited benefit, reducing storage and maintenance overhead</LI></UL><P>Each recommendation comes with a ready-to-execute SQL statement, so implementation is a copy-paste away.</P><H4 id="toc-hId-307700000">Key benefits</H4><UL><LI>Turns workload statistics into concrete index changes without requiring manual query analysis</LI><LI>Reduces both performance gaps (missing indexes) and unnecessary overhead (unused indexes)</LI><LI>Eliminates the guesswork of index selection by grounding recommendations in actual query behavior</LI></UL><DIV class=""><H4 id="toc-hId--386530600">Getting started</H4>Enable from the Advisors tab. Requires <CODE>CATALOG READ</CODE> and the relevant <CODE>INDEX</CODE> privileges on the target schemas or tables.<DIV class=""><STRONG>Note:</STRONG> The current version supports filter queries on single tables. Join query support is on the roadmap.</DIV></DIV></DIV></DIV><!-- ── NSE Advisor ── --><DIV class=""><DIV class=""><H3 id="toc-hId--289641098"><FONT color="#3366FF">NSE (Native Storage Extension) Advisor</FONT></H3></DIV><DIV class=""><H4 id="toc-hId--779557610">The problem it solves</H4><P>In-memory storage is SAP HANA Cloud's greatest performance asset - and its most expensive one. Not all data deserves to live in memory. Large, infrequently accessed tables or partitions silently inflate your memory footprint and cost, while small, hot objects may sit on disk unnecessarily.</P><H4 id="toc-hId--976071115">What the advisor does</H4><P>It collects access statistics over a configurable window (up to 90 days) and calculates scan density per object - tables, partitions, and columns - to determine data "temperature." Based on configurable thresholds, it generates two types of recommendations:</P><UL><LI><STRONG>Cost recommendations:</STRONG> Move an object to page-loadable (disk) storage to reduce memory consumption</LI><LI><STRONG>Performance recommendations:</STRONG> Move an object to column-loadable (memory) storage to improve response times</LI></UL><P>Each recommendation includes a confidence score reflecting how representative the observed workload window is.</P><H4 id="toc-hId--1172584620">Key benefits</H4><UL><LI>Reduces in-memory footprint without degrading performance for hot data</LI><LI>Surfaces cost savings opportunities that are invisible without systematic access pattern analysis</LI><LI>Confidence scores let you prioritize high-certainty recommendations and defer uncertain ones</LI><LI>Supports bulk operations so you can apply changes incrementally and at scale</LI></UL><DIV class=""><H4 id="toc-hId--1369098125">Getting started</H4>Click NSE Advisor in Recommendation Card and enable the NSE Advisor in NSE Recommendations tab. Configure your collection window, hot/cold object percentage targets, and minimum object size thresholds. Requires <CODE>TABLE ADMIN</CODE>, <CODE>PARTITION ADMIN</CODE>, and <CODE>CATALOG READ</CODE> system privileges.</DIV></DIV></DIV><!-- ── ECN Advisor ── --><DIV class=""><DIV class=""><H3 id="toc-hId--1272208623"><FONT color="#3366FF">ECN (Elastic Compute Node) Advisor</FONT></H3></DIV><DIV class=""><H4 id="toc-hId--1762125135">The problem it solves</H4><P>Some workloads are predictably periodic - end-of-month reporting, nightly batch jobs, seasonal peaks. Provisioning a coordinator large enough to handle the peak means over-paying for compute during the other 90% of the time. But manually identifying which queries to route, when to spin up an ECN, and when to tear it down is operationally complex.</P><H4 id="toc-hId--1958638640">What the advisor does</H4><P>It analyzes past coordinator workload against configurable memory and compute thresholds to identify peak windows where query routing to a temporary elastic compute node would bring the coordinator back within threshold. The resulting recommendation specifies the ECN size (memory and compute), the workload classes to route, and a provisioning/de-provisioning schedule.</P><P>You can apply the recommendation in two ways:</P><UL><LI><STRONG>ECN Policy (recommended):</STRONG> Creates an automated policy that provisions and de-provisions ECNs on your behalf - either on a fixed schedule (daily, weekly, monthly) or dynamically in response to live system load. The policy manages workload class routing automatically.</LI><LI><STRONG>Manual:</STRONG> Add the ECN via Manage Configuration, update workload class routing locations, and remove the ECN when the time range ends.</LI></UL><H4 id="toc-hId-2139815151">Key benefits</H4><UL><LI>Eliminates the need to maintain an oversized coordinator for occasional peaks</LI><LI>Right-sizing the coordinator for regular workload after implementing ECN routing can meaningfully reduce total cost of ownership</LI><LI>ECN Policies fully automate the provisioning lifecycle, removing operational overhead</LI></UL><DIV class=""><H4 id="toc-hId-2111485337">Getting started</H4>The ECN Advisor requires at least 5 vCPUs and at least one existing workload class. Enable from the Advisors tab, set memory and compute thresholds, define an analysis timeframe (30 minutes to 24 hours), and generate a recommendation. Enabling expensive statement tracing during the analysis window improves recommendation accuracy.</DIV></DIV></DIV><!-- ── Performance Class Advisor ── --><DIV class=""><DIV class=""><H3 id="toc-hId--2086592457"><FONT color="#3366FF">Performance Class Advisor</FONT></H3></DIV><DIV class=""><H4 id="toc-hId-1718458327">The problem it solves</H4><P>Knowing when your instance needs to be resized - and to which performance class - is rarely obvious until performance has already degraded. Reactive resizing means business continuity risk. Proactive resizing based on gut feel means unnecessary cost.</P><H4 id="toc-hId-1521944822">What the advisor does</H4><P>It runs automatically every Sunday, analyzing 2, 4, or 6 weeks of historical compute and memory utilization to determine whether your current performance class and instance size are still appropriate. When it identifies a mismatch, it produces a specific recommendation - which performance class and size to move to, and why.</P><H4 id="toc-hId-1325431317">Key benefits</H4><UL><LI>Automatically generated - switch it on once, and it runs weekly without any manual trigger</LI><LI>Grounds sizing decisions in actual historical workload rather than estimates or guesswork</LI><LI>Proactively surfaces the need to upsize before performance degradation reaches users</LI></UL><DIV class=""><H4 id="toc-hId-1128917812">Getting started</H4>Enable from the Advisors tab and choose your analysis window (2, 4, or 6 weeks). No additional configuration required.<DIV class=""><STRONG>Note:</STRONG> The current version supports up-sizing scenarios only. Down-sizing recommendations are on the roadmap.</DIV></DIV></DIV></DIV><!-- Getting started steps --><H2 id="toc-hId-1519210321">&nbsp;</H2><H2 id="toc-hId-1322696816"><FONT color="#008000">Getting Started with All Advisors</FONT></H2><P>All advisors are accessible from the <STRONG>Recommendations card</STRONG> on any instance's Overview page in SAP HANA Cloud Central:</P><OL class=""><LI>Open your instance's <STRONG>Overview</STRONG> page</LI><LI>Select the <STRONG>Recommendations</STRONG> card</LI><LI>Go to the <STRONG>Advisors</STRONG> tab (*NSE Advisor is located in NSE App as of today. It will be fully integrated into Recommendation App soon)</LI><LI>Select any advisor row to view details, enable it, and configure thresholds</LI></OL><P>You'll need the <STRONG>SAP HANA Cloud Administrator</STRONG> role collection to enable and configure advisors, and <STRONG>SAP HANA Cloud Viewer</STRONG> to read recommendations.</P><!-- What's coming --><DIV class=""><H3 id="toc-hId-832780304">&nbsp;</H3><H3 id="toc-hId-636266799"><FONT color="#666699">What's Coming — Analytics Advisor (Q3 2026)</FONT></H3><P>The advisor portfolio is actively expanding. <STRONG>Analytics Advisor</STRONG> will analyze OLAP query performance across your instance, identify slow calculation views, and recommend creating MDS cubes - materialized aggregation structures that dramatically reduce response times for repeated queries.</P><P>Database sizing, physical design, storage tiering, and compute management are all moving from reactive, expert-driven tasks to automated, data-driven operations. Each new advisor is a step toward a database that tunes itself.</P></DIV></DIV> 2026-07-02T13:05:27.997000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/what-s-new-in-sap-hana-cloud-july-2026/ba-p/14419199 What’s New in SAP HANA Cloud – July 2026 2026-07-03T11:29:14.472000+02:00 thomashammer https://community.sap.com/t5/user/viewprofilepage/user-id/122781 <P>The Q2 2026 release of SAP HANA Cloud continues to advance our vision of delivering an intelligent, AI-native, and enterprise-scale data platform. This quarter introduces innovations that make it easier to build AI-powered applications, manage large-scale enterprise workloads, simplify database administration, and strengthen security across your data landscape.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="blog.png" style="width: 725px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428619i31CE476C9DFE786B/image-dimensions/725x281?v=v2" width="725" height="281" role="button" title="blog.png" alt="blog.png" /></span></P><P>From deeper integration with SAP AI Core and new machine learning capabilities, to larger database instance sizes, enhanced data integration, and intelligent administration tools, this release allows you to unlock greater value from your data while reducing operational complexity.</P><P>Whether you're building next-generation AI applications, modernizing your enterprise data landscape, or optimizing business-critical workloads, SAP HANA Cloud continues to provide the foundation for secure, scalable, and intelligent data management.</P><P>Let's take a look at some of the key innovations introduced with the Q2 2026 release.</P><H2 id="toc-hId-1817435640">&nbsp;</H2><H2 id="toc-hId-1620922135">Agentic Data Access in SAP HANA Cloud</H2><P>We're taking another major step toward an AI-native data platform with new capabilities that enable agentic data access in SAP HANA Cloud.</P><P>At the core of this innovation is the ability to automatically generate a custom database objects knowledge graph. SAP HANA Cloud scans your database metadata including: tables, views, columns, and their relationships, and automatically builds a semantic representation of your data landscape.</P><P>By eliminating the need for manual semantic modeling, this knowledge graph provides a rich foundation for data discovery, reasoning, and AI-driven interactions. It enables both users and AI agents to better understand how enterprise data is organized and connected.</P><P>Building on this semantic foundation, we are introducing the Database Object Discovery Tool and the Data Retrieval Tool, available beginning in early July.</P><P>Today, working with enterprise data often requires you to understand complex database schemas, identify the right database objects, and manually write SQL queries to retrieve information. With these new capabilities, you can simply describe what they are looking for using natural language.</P><P>The Database Object Discovery Tool identifies the relevant database objects and relationships by leveraging the database objects knowledge graph. The Data Retrieval Tool then translates your request into an executable SQL statement and runs it directly within SAP HANA Cloud.</P><P>Together, these capabilities dramatically simplify the way we can discover, understand, and query enterprise data, without requiring deep SQL expertise or detailed knowledge of the underlying database schema.</P><P>Built on the multi-model foundation and Knowledge Graph Engine of SAP HANA Cloud, these innovations lay the groundwork for a new generation of conversational, agent-driven, and AI-powered data experiences, making it easier than ever to build intelligent applications on enterprise data.</P><P>To learn more about the agentic database tools, take a look at Shabana's blog <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/sap-hana-cloud-becomes-agentic-introducing-discovery-agent-amp-data-agent/ba-p/14257394" target="_self">post</A>.</P><P>&nbsp;</P><H2 id="toc-hId-1424408630">Natural Language Processing (NLP),&nbsp;Machine Learning and&nbsp;AI Core integration&nbsp;</H2><P><SPAN>With the Q2&nbsp;2026 release of SAP HANA Cloud, we continue to strengthen embedded&nbsp;NLP and&nbsp;Machine Learning capabilities&nbsp;with broader SAP AI Core integration, unlocking&nbsp;access to&nbsp;SAP RPT-1&nbsp;and&nbsp; LLMs&nbsp;and thus completely new type of AI-driven database applications.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Key highlights&nbsp;as part of the Predictive Analysis Library (PAL)&nbsp;include:</SPAN><SPAN>&nbsp;</SPAN></P><UL><LI><SPAN>A new&nbsp;function for synthetic data generation using&nbsp;<EM>diffusion models</EM>,&nbsp;e.g&nbsp;for securing data privacy when training custom AI models</SPAN><SPAN>&nbsp;</SPAN></LI><LI><SPAN>Time Series Forecasting function enhancements like support for&nbsp;<EM>lagged regressors</EM>&nbsp;with&nbsp;<EM>Additive Model&nbsp;Time Series Analysis</EM>&nbsp;(aka prophet) and easier time series cross validation using the&nbsp;<EM>Unified&nbsp;Time&nbsp;Series</EM>&nbsp;procedures for better forecast predictions</SPAN><SPAN>&nbsp;</SPAN></LI><LI><SPAN>Further enhancements include&nbsp;columnar analysis of large data sets&nbsp;using&nbsp;Online Univariate Analysis</SPAN><SPAN>&nbsp;</SPAN></LI></UL><P><SPAN>The full list of PAL&nbsp;</SPAN><SPAN>enhancements&nbsp;for&nbsp;SAP HANA Cloud 2026 Q2&nbsp;is documented&nbsp;</SPAN><A href="https://help.sap.com/whats-new/2495b34492334456a49084831c2bea4e?Category=Predictive+Analysis+Library&amp;Valid_as_Of=2026-06-01:2026-06-30&amp;locale=en-US" target="_blank" rel="noopener noreferrer"><SPAN>here.</SPAN></A>&nbsp;Additionally, th<SPAN>e python Machine Learning client for SAP HANA 2.29 is supporting several new functions&nbsp;(see&nbsp;</SPAN><A href="https://help.sap.com/doc/1d0ebfe5e8dd44d09606814d83308d4b/Latest/en-US/change_log.html" target="_blank" rel="noopener noreferrer"><SPAN>here</SPAN></A><SPAN>)&nbsp;like&nbsp;PCA-based vector dimension reduction or the text&nbsp;tf-idf&nbsp;operator in AutoML.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Key highlights for SAP AI Core integration include:</SPAN><SPAN>&nbsp;</SPAN></P><UL><LI><SPAN>Remote sources to SAP AI Core now supporting use of LLMs and SAP RPT-1&nbsp;as standardized and secure gateway for 3rd party LLMs access</SPAN><SPAN>&nbsp;</SPAN></LI><LI><SPAN>Benefit from on-the-fly predictions, without the need to create, train and manage use-case specific ML models&nbsp;by&nbsp;s</SPAN><SPAN>imply&nbsp;utilizing&nbsp;the&nbsp;new&nbsp;AI_TABULAR_PREDICTION SQL procedure to access tabular AI models&nbsp;like SAP RPT-1&nbsp;</SPAN><SPAN>&nbsp;</SPAN></LI><LI><SPAN>Ability to leverage LLM capabilities, directly from the database SQL processing engine, by s</SPAN><SPAN>imply&nbsp;utilizing&nbsp;the new AI_TEXT_COMPLETION SQL function to execute one-shot-tasks like text generation or text summarization based on your text data managed in SAP HANA Cloud, unlocking for completely new type of LLM-driven application scenarios.</SPAN><SPAN>&nbsp;</SPAN></LI></UL><P><SPAN>These enhancements focus on making tabular AI predictions using custom AI </SPAN><SPAN>models or semantic data retrieval scenario easier to use and applied to&nbsp;real business&nbsp;application data.</SPAN><SPAN>&nbsp;</SPAN></P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="thomashammer_0-1783013936469.png" style="width: 795px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428611iC30E6D7085D6554B/image-dimensions/795x306?v=v2" width="795" height="306" role="button" title="thomashammer_0-1783013936469.png" alt="thomashammer_0-1783013936469.png" /></span></P><P>&nbsp;</P><P><SPAN>&nbsp;</SPAN></P><H2 id="toc-hId-1227895125">Right-sizing your instance with<SPAN><STRONG>&nbsp;Performance Class Advisor</STRONG></SPAN></H2><P><SPAN>Knowing whether your SAP HANA Cloud instance is the right size&nbsp;shouldn't&nbsp;require manual analysis or reactive firefighting. The Performance Class Advisor changes that by continuously&nbsp;monitoring&nbsp;your instance's&nbsp;compute&nbsp;and memory usage and automatically generating right-sizing recommendations every Sunday, without any manual trigger.</SPAN></P><P><SPAN>You choose a lookback window of 2, 4, or&nbsp;6 weeks&nbsp;to match your workload patterns, and the advisor checks whether CPU or memory usage has exceeded 80% during that period, if either threshold is breached, a recommendation is triggered.&nbsp;</SPAN><SPAN>The advisor then calculates the&nbsp;optimal&nbsp;upsized spec, targeting approximately 20%&nbsp;additional&nbsp;capacity and snapping to the nearest valid instance configuration.</SPAN></P><P><SPAN>In this&nbsp;initial&nbsp;release, recommendations are limited to upsizing only, as downsizing carries higher risk given that short lookback windows may not capture infrequent but critical workload spikes, and that is in our backlog to investigate further. The benefits are threefold: right-sizing guidance based on actual historical workload patterns rather than guesswork; proactive resource management that surfaces capacity pressure before it impacts business continuity, giving you time to act rather than react; and automated weekly insights with zero manual effort — switch it on once, and stay informed from there.</SPAN><SPAN>&nbsp;</SPAN></P><P>Learn more on SAP HANA Cloud's Intelligent Advisors in <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/stop-tuning-your-database-manually-let-sap-hana-cloud-s-intelligent/ba-p/14422705" target="_blank">this blogpost</A> by&nbsp;<a href="https://community.sap.com/t5/user/viewprofilepage/user-id/159990">@jinheejeong</a>&nbsp;</P><P>&nbsp;</P><H2 id="toc-hId-1031381620"><SPAN><STRONG>Automated SQL Plan Management with SQL Plan Advisor</STRONG></SPAN><SPAN>&nbsp;</SPAN></H2><P><SPAN>SQL execution performance can fluctuate for&nbsp;various reasons. A common culprit is a change in the execution plan within the SQL Plan Cache, either because a cached plan becomes suboptimal over time, or because an inefficient plan was cached based on the&nbsp;initial&nbsp;parameter sets used during compilation.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>To better manage these regressions,&nbsp;<STRONG>SQL Plan Advisor</STRONG>&nbsp;was introduced in QRC 2/2024. It automatically detects degraded query executions by analyzing historical performance statistics stored via&nbsp;<STRONG>SQL Plan Stability</STRONG>. To resolve a performance dip, older, historical plans had to be verified to ensure they performed faster than the currently cached plan. However, this verification and testing process could only be triggered manually.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>With QRC 2/2026,&nbsp;<STRONG>this testing process is now fully automated</STRONG>. The system autonomously executes faster historical plans to verify their effectiveness, swapping out the degraded cached plan to seamlessly restore&nbsp;optimal&nbsp;database performance.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Please check out the details in the blog by&nbsp;<a href="https://community.sap.com/t5/user/viewprofilepage/user-id/202976">@Taesuk</a>&nbsp; found&nbsp;</SPAN><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/automated-sql-plan-management-to-recover-from-degraded-performance/ba-p/14419577" target="_blank"><SPAN>here</SPAN></A>.<SPAN>&nbsp;</SPAN></P><P>&nbsp;</P><H2 id="toc-hId-834868115">Database Administration &amp; Developer Productivity</H2><H3 id="toc-hId-767437329">Import &amp; Export&nbsp;&nbsp;</H3><P><SPAN>SAP HANA Cloud Central&nbsp;will&nbsp;soon&nbsp;be adding a&nbsp;new top-level&nbsp;application&nbsp;named&nbsp;Import and Export designed&nbsp;to enable users to quickly complete import and export operations within SAP HANA Cloud. This application provides guided wizards for importing data into a table, exporting data from a table or view, and importing and exporting both catalog objects and SAP HANA deployment infrastructure containers.&nbsp;&nbsp;Supported storage providers include SAP HANA Cloud data lake, Amazon S3, Microsoft Azure storage, and Google Cloud storage.&nbsp;Users can also view the generated SQL for any import or export task directly within the wizard.</SPAN><SPAN>&nbsp;</SPAN></P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="thomashammer_3-1783013999341.png" style="width: 631px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428614i7B165F92C3CFD5B5/image-dimensions/631x369?v=v2" width="631" height="369" role="button" title="thomashammer_3-1783013999341.png" alt="thomashammer_3-1783013999341.png" /></span></P><P><SPAN>&nbsp;</SPAN><SPAN>Additionally,&nbsp;the export&nbsp;wizard will support the ability to provide a query to further filter the data being exported.</SPAN><SPAN>&nbsp;</SPAN></P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="thomashammer_2-1783013991448.png" style="width: 632px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428613i861F63F890C8309C/image-dimensions/632x295?v=v2" width="632" height="295" role="button" title="thomashammer_2-1783013991448.png" alt="thomashammer_2-1783013991448.png" /></span></P><P>&nbsp;</P><H2 id="toc-hId-441841105">SQL Console&nbsp;</H2><P><U>Improved Tab Management&nbsp;</U><BR /><SPAN>Previously if a&nbsp;SQL console tab&nbsp;was&nbsp;closed&nbsp;accidently,&nbsp;its contents were gone.&nbsp;Users can now restore closed tabs&nbsp;using&nbsp;the&nbsp;<STRONG>Restore Last Closed Tab</STRONG>&nbsp;option in the&nbsp;toolbar.&nbsp;Additionally, tabs can be renamed and reordered, making it easier to stay organized when working across multiple queries at once.&nbsp;</SPAN><SPAN>&nbsp;</SPAN></P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="thomashammer_4-1783014025884.png" style="width: 598px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428615i7450DD56208D68B4/image-dimensions/598x263?v=v2" width="598" height="263" role="button" title="thomashammer_4-1783014025884.png" alt="thomashammer_4-1783014025884.png" /></span></P><P><SPAN>A new search capability is also available&nbsp;from the actions menu, that lets users search across the content of all open tabs&nbsp;enabling a&nbsp;specific query&nbsp;to be&nbsp;quickly&nbsp;located.&nbsp;</SPAN><SPAN>&nbsp;</SPAN></P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="thomashammer_5-1783014033226.png" style="width: 571px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428616iCC834D7688EF5E5B/image-dimensions/571x314?v=v2" width="571" height="314" role="button" title="thomashammer_5-1783014033226.png" alt="thomashammer_5-1783014033226.png" /></span></P><P><U>Enhanced Syntax Help Panel<STRONG>&nbsp;</STRONG>&nbsp;</U><BR /><SPAN>The syntax help panel has also received a refresh. Object details are now collapsible, reducing visual clutter when users only need a quick reference.&nbsp;Columns and parameters are also&nbsp;searchable, which allows users to find what they need without scrolling through extensive lists.&nbsp;These columns are now clickable as well – selecting one directly inserts it into the editor.&nbsp;</SPAN><SPAN>&nbsp;</SPAN></P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="thomashammer_6-1783014051187.png" style="width: 587px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428617i2BEB452B4DCF74A6/image-dimensions/587x418?v=v2" width="587" height="418" role="button" title="thomashammer_6-1783014051187.png" alt="thomashammer_6-1783014051187.png" /></span></P><P><U>SQL Results<STRONG>&nbsp;</STRONG>&nbsp;</U><BR /><SPAN>Users now also have more control within SQL results, which has been enhanced with options to handle how data is displayed and refreshed.&nbsp;Users can customize which columns they want to be&nbsp;shown&nbsp;or&nbsp;hidden&nbsp;and&nbsp;also&nbsp;reorder them in whichever way that suits their workflow.&nbsp;Results can also be re-executed on&nbsp;demand&nbsp;without requiring parameters to be re-entered.&nbsp;</SPAN><SPAN>&nbsp;</SPAN></P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="thomashammer_0-1783014096304.png" style="width: 601px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428618i280B7ACAFC93F0DB/image-dimensions/601x314?v=v2" width="601" height="314" role="button" title="thomashammer_0-1783014096304.png" alt="thomashammer_0-1783014096304.png" /></span></P><P>&nbsp;</P><H2 id="toc-hId-245327600"><SPAN><STRONG>Data Integration</STRONG></SPAN><SPAN>&nbsp;(Federation &amp; Replication)</SPAN></H2><H3 id="toc-hId-177896814"><SPAN><STRONG>SQL on Files: Apache Iceberg REST Catalog support in SAP HANA Cloud</STRONG></SPAN><SPAN>&nbsp;</SPAN></H3><P><SPAN>Direct read-only access to Apache Iceberg tables in object storage, which has been introduced since QRC 3/2025, has offered real-time analytics on Iceberg&nbsp;lakehouses&nbsp;without moving data into SAP HANA Cloud.&nbsp;With QRC 2/2026, it evolves further by supporting&nbsp;<STRONG>Apache Iceberg REST Catalogs</STRONG>,&nbsp;which is&nbsp;how&nbsp;the&nbsp;modern enterprise&nbsp;lakehouses&nbsp;centralize table management, access control, and metadata governance.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Through the&nbsp;new&nbsp;<EM>icebergcatalog</EM>&nbsp;adapter, SAP HANA Cloud now connects directly to spec-compliant Iceberg REST Catalogs across AWS, Azure, and GCP, bringing enterprise-grade integration to your Iceberg&nbsp;lakehouse&nbsp;with minimal effort.&nbsp;Note&nbsp;that while the adapter is designed to work with any spec-compliant&nbsp;Iceberg&nbsp;REST Catalog, only a defined set of providers is&nbsp;officially certified by SAP per QRC.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>To learn more about SAP HANA Cloud's new Apache Iceberg REST Catalog support, take a look at <a href="https://community.sap.com/t5/user/viewprofilepage/user-id/204092">@SeungjoonLee</a>'s blog post&nbsp;</SPAN><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/going-beyond-the-tip-of-the-iceberg-with-sap-hana-cloud-sql-on-files/ba-p/14413871" target="_blank"><SPAN>here</SPAN></A><SPAN>.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN><STRONG>S/4HANA CDS View Entity replication to SAP HANA Cloud</STRONG></SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>S/4HANA Federated access was introduced in QRC 1/2024, as detailed in the blog&nbsp;</SPAN><A href="https://community.sap.com/t5/technology-blogs-by-sap/taking-data-federation-to-the-next-level-accessing-remote-abap-cds-view/ba-p/13635034" target="_blank"><SPAN>Taking Data Federation to the Next Level: Accessing Remote ABAP CDS View Entities in SAP HANA Cloud.</SPAN></A><SPAN>&nbsp;Building on this, QRC 2/2026 introduces support for replicating data directly to SAP HANA Cloud.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>CDS view replication allows you to copy data from SAP S/4HANA CDS&nbsp;view&nbsp;entities into local tables within SAP HANA Cloud, enabling fast local analytics without the need for remote federated queries.<BR /><BR />Key benefits include:</SPAN><SPAN>&nbsp;</SPAN></P><UL><LI><SPAN><STRONG>Lower Latency:</STRONG></SPAN><SPAN>&nbsp;Data&nbsp;resides&nbsp;locally, and efficiency is maximized by replicating only changed rows.</SPAN><SPAN>&nbsp;</SPAN></LI><LI><SPAN><STRONG>Reduced Source Overhead:</STRONG></SPAN><SPAN>&nbsp;Impact on the source system is minimized because replication can be scheduled at custom intervals or run entirely on demand.</SPAN><SPAN>&nbsp;</SPAN></LI><LI><SPAN><STRONG>Flexible Filtering:</STRONG></SPAN><SPAN>&nbsp;Data can be filtered by specific rows, selected columns, or both.</SPAN><SPAN>&nbsp;</SPAN></LI></UL><P><SPAN>You can configure replication via SQL or by using the Data Replication tool in SAP HANA Cloud Central.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>To learn more about SAP HANA Cloud’s new S/4HANA CDS View replication, check out <a href="https://community.sap.com/t5/user/viewprofilepage/user-id/202976">@Taesuk</a>&nbsp;’s blog post&nbsp;</SPAN><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/replicating-sap-s-4hana-cds-view-entities-into-sap-hana-cloud-a-practical/ba-p/14419559" target="_blank"><SPAN>here</SPAN></A><SPAN>.</SPAN><SPAN>&nbsp;</SPAN></P><P>&nbsp;</P><H2 id="toc-hId-199554947"><SPAN>Security, Data Protection &amp; Privacy<BR /><BR /></SPAN></H2><H2 id="toc-hId-3041442"><SPAN><STRONG>Post-quantum cryptography (PQC) for Transport Layer Security (TLS) connections</STRONG></SPAN><SPAN><STRONG>&nbsp;</STRONG></SPAN></H2><P><SPAN>With the Q2 2026 release, SAP HANA Cloud&nbsp;database&nbsp;introduces post-quantum cryptography (PQC) for TLS connections, providing quantum-resistant protection for all data in transit between clients and SAP HANA Cloud databases.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Post-quantum cryptography addresses the "harvest now, decrypt later" threat — where adversaries capture encrypted traffic today with the intent of decrypting it once quantum computers become capable. By enabling PQC now, SAP HANA Cloud ensures that data encrypted today&nbsp;remains&nbsp;protected against future quantum attacks.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Key capabilities include:</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN><STRONG>Hybrid Post-Quantum Key Exchange</STRONG></SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>A transparent ML-KEM key establishment negotiated during TLS 1.3 handshakes, combining classical and quantum-resistant algorithms. Enabled by default for all SAP HANA Cloud databases with no customer configuration&nbsp;required.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN><STRONG>Full Backward Compatibility</STRONG></SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Clients with PQC support (SAP Client using&nbsp;OpenSSL 3.5+, SAP HANA client 2.29+) automatically negotiate the quantum-resistant hybrid exchange. Clients without PQC support safely fall back to classical key exchange with zero disruption to existing applications.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN><STRONG>Built-in Observability</STRONG></SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>New SSL_KEY_EXCHANGE and SSL_SIGNATURE columns in M_CONNECTIONS and M_OUTBOUND_NETWORK_IO system views, enabling administrators to verify PQC adoption per connection and&nbsp;identify&nbsp;clients that need upgrading.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>The feature is aligned with NIST post-quantum standards (ML-KEM / FIPS 203) ahead of&nbsp;anticipated&nbsp;regulatory mandates and requires no customer action to activate.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Typical use cases include verifying quantum-safe connection status across the landscape&nbsp;and meeting compliance requirements for post-quantum readiness.</SPAN><SPAN>&nbsp;</SPAN></P><P>&nbsp;</P><H2 id="toc-hId--193472063">Access Logging &amp; Column-Level Security for SQL Views</H2><P>SAP HANA Cloud introduces two complementary capabilities for protecting sensitive data: access logging for columns&nbsp;containing&nbsp;sensitive data and column-level security for SQL views. Together, these address&nbsp;two&nbsp;sides of the compliance equation - preventing unauthorized access and detecting all access for audit purposes - enabling compliant direct SQL access to data&nbsp;containing&nbsp;sensitive columns.&nbsp;</P><H3 id="toc-hId--683388575"><STRONG>The Context</STRONG>&nbsp;</H3><P>Regulations like GDPR require both prevention of unauthorized access to personal data AND demonstrable accountability for all access. Direct SQL access to HANA databases&nbsp;containing&nbsp;sensitive data requires enforcement and audit trails to meet these compliance requirements.&nbsp;</P><P>These two capabilities address this:&nbsp;</P><TABLE width="0"><TBODY><TR><TD width="75"><P>Prevent:&nbsp;</P></TD><TD width="206"><P><STRONG>Column-Level Security for SQL Views</STRONG>&nbsp;</P></TD><TD width="348"><P>Ensures unauthorized users cannot see sensitive columns (returned as NULL)&nbsp;</P></TD></TR><TR><TD width="75"><P>Detect:&nbsp;</P></TD><TD width="206"><P><STRONG>Access Logging for Sensitive Columns</STRONG>&nbsp;</P></TD><TD width="348"><P>Logs who accessed which sensitive columns, when, and on which table (values never logged)&nbsp;</P></TD></TR></TBODY></TABLE><H3 id="toc-hId--879902080">Access Logging&nbsp;</H3><P>Access logging is activated directly in the table definition using the AUDIT (READ | NONE) syntax. Once columns are marked, read access is automatically reported via the predefined "SAP - column access" audit policy, as no custom policy creation is needed. Both direct and indirect access is reported, including access via views or during procedure execution. The audit log entry includes the schema name, table name, and the list of accessed columns (quoted and comma-separated), while the actual data values and SQL statements are never logged&nbsp;-&nbsp;ensuring the audit trail provides accountability without creating a secondary repository of sensitive data.&nbsp;</P><H3 id="toc-hId--1076415585">Column-Level Security&nbsp;</H3><P>Column-level security for SQL views&nbsp;provides&nbsp;a dynamic access control mechanism. Columns on SQL views are tagged with DPP classification&nbsp;identifying&nbsp;Sensitive Personal Data, and the authorization model&nbsp;determines&nbsp;which tagged columns are&nbsp;permitted&nbsp;for access.&nbsp;At query&nbsp;execution time,&nbsp;the engine evaluates&nbsp;authorization&nbsp;dynamically – this evaluation is not tied to the current database user but can rely on any information in the session context&nbsp;or other tables.&nbsp;Authorized columns return actual values, while unauthorized columns return NULL. The key design decision is that columns are masked, not omitted - query structure&nbsp;remains&nbsp;stable for all consumers, meaning downstream dashboards, pipelines, and applications continue to function without modification.&nbsp;</P><H3 id="toc-hId--1272929090"><STRONG>Combined Value</STRONG>&nbsp;</H3><P>Together, these capabilities enable compliant direct SQL access to sensitive data in SAP HANA Cloud. Column-level security ensures users see only data they&nbsp;are authorized to&nbsp;access, while access logging provides auditable evidence of all access events, satisfying both the prevention and accountability requirements of GDPR and similar regulations.<BR />The operational overhead is minimal: DDL-based activation for access logging, a system-managed audit policy, and dynamic authorization evaluation for column-level security&nbsp;-&nbsp;with no duplicate views, no application-layer workarounds, and no custom configuration frameworks&nbsp;required.&nbsp;</P><P>&nbsp;</P><H2 id="toc-hId--1176039588"><SPAN>Increase in-memory capacity with up to 16 TB and 24 TB Instances on AWS</SPAN><SPAN>&nbsp;</SPAN></H2><P>As enterprise workloads continue to grow in size and complexity, organizations increasingly require larger in-memory database configurations to support mission-critical transactional and analytical applications.<BR />With the Q2 2026 release, SAP HANA Cloud expands its scalability by introducing support for <STRONG>16 TB and 24 TB database instances on AWS</STRONG>.</P><P>These new instance sizes are designed for customers running some of the most demanding enterprise workloads, providing the processing power and memory capacity required for large-scale SAP and custom applications while maintaining the simplicity of a fully managed cloud service.</P><P>To ensure predictable performance and guaranteed resource availability, these instances are deployed on <STRONG>dedicated, pre-reserved physical infrastructure</STRONG> that is allocated exclusively to a customer's subaccount within a selected AWS availability zone. The reserved capacity remains available throughout the three-year reservation period, regardless of whether the database instance is running, stopped, paused, or temporarily deprovisioned.</P><P>Customers can request these instance sizes directly through SAP HANA Cloud Central by opening an SAP support ticket. Once the requested capacity has been approved and the dedicated infrastructure has been provisioned, the new database instance can be deployed against the reserved resources.</P><P>Support for 16 TB and 24 TB instances is currently available in selected AWS regions. For the latest information on supported regions and availability zones, please refer to the SAP HANA Cloud <A href="https://help.sap.com/docs/hana-cloud/sap-hana-cloud-administration-guide/large-sap-hana-database-instances" target="_blank" rel="noopener noreferrer">documentation</A>.</P><P>&nbsp;</P><H2 id="toc-hId--1372553093"><SPAN><STRONG>Data Lake Innovations</STRONG></SPAN></H2><H3 id="toc-hId--1862469605"><SPAN><STRONG>Thin Clone Decoupling for Data Lake Relational Engine</STRONG></SPAN><SPAN>&nbsp;</SPAN></H3><P><SPAN>Building on the thin clone capability in SAP HANA Cloud, data lake relational engine, which is designed for creation and deployment of clones of petabyte-scale databases in development and testing environments in just a few minutes, we are introducing&nbsp;</SPAN><SPAN>thin clone decoupling.&nbsp;</SPAN><SPAN>Thin clone decoupling promotes a thin clone to a fully standalone SAP HANA Cloud, data lake relational engine instance. Once the decoupling process completes, the instance&nbsp;operates&nbsp;independently, completely detached from the parent instance.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN>Upon decoupling, backups are enabled on the decoupled instance, making it eligible for disaster recovery scenarios. The decoupled instance can itself serve as a source for creating new thin clones.</SPAN><SPAN>&nbsp;</SPAN></P><P><SPAN> </SPAN><SPAN>Thin clone decoupling opens a clear path for promoting development or test databases to production-grade instances improving data reliability and integrity.</SPAN><SPAN>&nbsp;</SPAN></P><P>&nbsp;</P><H2 id="toc-hId--1597396412"><SPAN><STRONG>Innovations on Calculation View features</STRONG></SPAN></H2><P>The <A class="" title="https://community.sap.com/t5/technology-blog-posts-by-sap/calculation-view-features-of-2026-qrc2/ba-p/14415862" href="https://community.sap.com/t5/technology-blog-posts-by-sap/calculation-view-features-of-2026-qrc2/ba-p/14415862" target="_blank">calculation view features blog for 2026 QRC2</A>&nbsp;by <a href="https://community.sap.com/t5/user/viewprofilepage/user-id/239612">@jan_zwickel</a>&nbsp;covers several new calculation view capabilities. Two complementary deployment options help to reduce deployment time and protect privileges: <STRONG>optimized redeployment</STRONG> skips unchanged dependent views instead of dropping and recreating them all, and the new <STRONG>replace option</STRONG> performs in-place replacement so that cross-container privileges are not lost.</P><P>On the scalability front, <STRONG>Active/Active read-enabled hints</STRONG> can be set at the individual calculation view level for fine-grained query routing to read replicas. A <STRONG>Join-to-Non-Equi-Join conversion</STRONG> in the context menu eliminates the need to rebuild join logic from scratch when switching from equi joins to non-equi joins. The support of <STRONG>calculated measures executed at query runtime&nbsp;</STRONG>allow defining runtime calculations for <STRONG>MDS Cubes</STRONG> in calculation views. Finally, a new <STRONG>status dialog</STRONG> surfaces the view's connection state, read-only reasons, data classification inconsistencies, and BDC integration status when a calculation view is opened.</P><P>&nbsp;</P><HR /><P>&nbsp;</P><P>Thanks for taking the time to explore the What’s New in SAP HANA Cloud on Q2 2026 innovations!</P><P>If you’d like to go deeper into the technical details of the Q2 2026 release, we recommend visiting the&nbsp;<A href="https://help.sap.com/whats-new/2495b34492334456a49084831c2bea4e" target="_blank" rel="noopener noreferrer">What’s New Viewer</A>&nbsp;in the technical documentation. It provides a comprehensive overview of the full release scope. To stay informed about upcoming innovations, announcements, and best practices, follow the&nbsp;<A href="https://community.sap.com/t5/c-khhcw49343/SAP+HANA+Cloud/pd-p/73554900100800002881" target="_blank">SAP HANA Cloud tag</A>.</P><P>You can also find our latest blog posts by searching for the&nbsp;<A href="https://community.sap.com/t5/tag/whatsnewinsaphanacloud/tg-p" target="_blank">#whatsnewinsaphanacloud</A>&nbsp;hashtag.</P><P>If you missed earlier What’s New webinars, you’ll find recordings of past sessions and upcoming events in our&nbsp;<A href="https://www.youtube.com/playlist?list=PL3ZRUb1AKkpTDZQgENtRcupp6vsNg8NHN" target="_blank" rel="noopener nofollow noreferrer">YouTube playlist</A>. Our What’s New webinar covering the Q2 2026 innovations will be posted on this playlist very soon. You can also download the slides from the webinar <A href="https://dam.sap.com/mac/u/a/kABQpTr?rc=10&amp;doi=SAP1324444" target="_blank" rel="noopener noreferrer">here</A>.</P><P>Do you have questions about SAP HANA Cloud or want to share your thoughts on the new features?<BR />Join the conversation in the&nbsp;<A href="https://community.sap.com/t5/technology-q-a/qa-p/technology-questions" target="_blank">SAP HANA Cloud Community Q&amp;A</A>&nbsp;or leave a comment below. We’d love to hear from you!</P> 2026-07-03T11:29:14.472000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-1-why-knowledge-graph/ba-p/14417381 Knowledge Graph Agent on SAP HANA Cloud Series – Part 1: Why Knowledge Graph for SAP Data 2026-07-04T15:44:34.848000+02:00 ClaudioJP https://community.sap.com/t5/user/viewprofilepage/user-id/1509109 <P class="">This blog is part of a blog series on building AI Agents with SAP BDC (Business Data Cloud) Data Products and SAP HANA Cloud Knowledge Graph Engine:</P><UL><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-1-why-knowledge-graph/ba-p/14417381" target="_self">Knowledge Graph Agent on SAP HANA Cloud Series – Part 1: Why Knowledge Graph for SAP Data</A></LI><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-2-1-ontology-standards/ba-p/14417385" target="_self">Knowledge Graph Agent on SAP HANA Cloud Series – Part 2-1: Ontology Standards and Concepts</A></LI><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-2-2-designing-the/ba-p/14274465" target="_self">Knowledge Graph Agent on SAP HANA Cloud Series – Part 2-2: Designing the Ontology for SAP Data</A>&nbsp;</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series – Part 3: Data Pipeline from BDC to HANA Cloud KGE</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series – Part 4: Building an Agent with SPARQL and SQL</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series&nbsp;– Part 5: Lessons Learned</LI></UL><P><SPAN>Enterprise systems are vast, with thousands of tables, hundreds of business processes, and countless APIs, all interconnected in ways only domain experts fully grasp. For AI Agents to be genuinely useful in this environment, they need to understand not just what the data is, but how it all connects: which Product belongs to which Sales Order, which WorkCenter is located in which Plant, which Supplier serves which material. This understanding does not come from the data itself. It needs a dedicated semantic layer that makes business meaning explicit and machine-readable. Knowledge Graphs provide exactly that layer. SAP HANA Cloud now offers a native Knowledge Graph Engine to build it on top of SAP BDC Data Products. This series walks through the full journey: from why this layer matters, to how to design and build it, to how an AI Agent uses it to answer real business questions.</SPAN></P><P class="">In this post, we describe why a Knowledge Graph is the right semantic foundation for AI Agents querying SAP business data, and how SAP BDC Data Products combined with HANA Cloud Knowledge Graph Engine make this practical.</P><P class=""><STRONG>Table of contents:</STRONG></P><OL class=""><LI>Introduction</LI><LI>Why Knowledge Graph now</LI><LI>The problem: enterprise questions are not single-table</LI><LI>Vector RAG and Knowledge Graph: different strengths</LI><LI>What a Knowledge Graph solves: relationships made explicit</LI><LI>Why SAP BDC Data Products</LI><LI>One Domain Model: an extensible vocabulary foundation</LI><LI>The division of labor: KG and SQL</LI><LI>The full journey: from Competency Question to AI Agent</LI></OL><HR /><H1 id="1-introduction" id="toc-hId-1688295222">1. Introduction</H1><P class="">Knowledge Graphs have a steep learning curve. Ontology, Triple Store, SPARQL, RDF, OWL, and SHACL introduce a large amount of unfamiliar terminology that can be overwhelming at first. Anyone comfortable with relational databases and SQL is likely to ask "How is this different from a SQL JOIN?" early on, and to question whether KG is necessary when Vector RAG is already available.</P><P class="">The good news is that the field already has well-organized foundational material:</P><UL class=""><LI>Stanford's<SPAN>&nbsp;</SPAN><EM>Ontology Development 101</EM><SPAN>&nbsp;</SPAN>is the standard introduction to ontology design, structured as a seven-step methodology.</LI><LI>The W3C<SPAN>&nbsp;</SPAN><EM>Direct Mapping</EM><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><EM>R2RML</EM><SPAN>&nbsp;</SPAN>recommendations define precise rules for translating relational data into RDF.</LI><LI>SHACL is the W3C standard for validating graph data integrity, and OWL provides the standard vocabulary for inference rules.</LI></UL><P class="">The challenge is that these resources are<SPAN>&nbsp;</SPAN><STRONG>scattered</STRONG>. Each is authoritative within its own scope, but none of them answers the practical question:<SPAN>&nbsp;</SPAN><EM>"To build an AI Agent on enterprise data, in what order should these pieces be combined, and how?"</EM><SPAN>&nbsp;</SPAN>That question can only be answered by building.</P><H3 id="what-this-series-aims-to-do" id="toc-hId-1749947155">What this series aims to do</H3><P class="">This series weaves those scattered standards together with SAP's strengths into a single narrative. While building an AI Agent on top of SAP BDC Data Products with the HANA Cloud Knowledge Graph Engine, we documented how the theory applies to real SAP data and where the standards stop short and require domain decisions.</P><P class="">For anyone heading down a similar path, this series should serve as a useful starting framework. With the standards already established by others as a foundation, you can focus on the SAP domain decisions that actually matter.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span><STRONG>A note on the workflow.</STRONG></P><P class="">In production, ontology design and long-term maintenance should rely on modeling and governance tools such as Protégé or Metaphactory. This series uses AI-assisted Vibe Coding as a pragmatic way to create a first draft quickly, especially for scenario-based projects, and to learn the domain decisions that must later be reviewed and refined in dedicated modeling tools.</P></BLOCKQUOTE><H3 id="how-to-follow-along" id="toc-hId-1553433650">How to follow along</H3><P class="">The actual production data cannot be shared publicly. However, all code and project structure used throughout this series is available in the GitHub repo below. No sample data is included, but you can use the same structure and code as a reference to apply the approach in your own SAP environment or a different domain.</P><P class=""><span class="lia-unicode-emoji" title=":file_folder:">📁</span><SPAN>&nbsp;</SPAN><A href="https://github.com/claudiopark86/sap-bdc-kge-agent-workshop" target="_blank" rel="noopener nofollow noreferrer">github.com/claudiopark86/sap-bdc-kge-agent-workshop</A></P><P class="">The final output this series builds, a KG-powered AI Agent for supply chain risk analysis, can be seen in action here before you begin:</P><P class=""><div class="video-embed-center video-embed"><iframe class="embedly-embed" src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2FAYj9MbEkspM%3Ffeature%3Doembed&amp;display_name=YouTube&amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DAYj9MbEkspM&amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2FAYj9MbEkspM%2Fhqdefault.jpg&amp;type=text%2Fhtml&amp;schema=youtube" width="200" height="113" scrolling="no" title="AI Agent on BDC data with KG and NL2SQL" frameborder="0" allow="autoplay; fullscreen; encrypted-media; picture-in-picture" allowfullscreen="true"></iframe></div></P><HR /><H1 id="2-why-knowledge-graph-now" id="toc-hId-1098754707">2. Why Knowledge Graph Now</H1><P class="">Building useful AI Agents has made one thing clearer than ever:<SPAN>&nbsp;</SPAN><STRONG>context matters.</STRONG></P><P class="">An Agent answering a user question needs to know<SPAN>&nbsp;</SPAN><EM>where to look</EM><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><EM>what is connected to what</EM>. Enterprise systems contain a vast number of business processes and data attributes, all entangled. An AI Agent needs to understand this landscape and produce accurate answers for mission-critical work running on top of it.</P><P class=""><STRONG>Can all of this context be injected through the prompt?</STRONG><SPAN>&nbsp;</SPAN>The scale is too large. Embedding every table, every relationship, and every business rule into a prompt is not realistic, and even if it were, there is no guarantee that the LLM would retrieve the right pieces from such a payload.</P><P class=""><STRONG>What about vector similarity search?</STRONG><SPAN>&nbsp;</SPAN>Vector RAG is strong at finding<SPAN>&nbsp;</SPAN><EM>"similar documents,"</EM><SPAN>&nbsp;</SPAN>but it has limitations when it comes to<SPAN>&nbsp;</SPAN><STRONG>structural relationship traversal</STRONG><SPAN>&nbsp;including</SPAN>&nbsp;questions like<SPAN>&nbsp;</SPAN><EM>"Which materials does this supplier provide, which finished goods do those materials feed into via the BOM, and which customers buy those finished goods?"</EM><SPAN>&nbsp;</SPAN>Similarity finds semantic neighbors; it does not follow an explicit relationship chain.</P><P class="">A Knowledge Graph addresses both limitations from a different angle. Relationships are stored explicitly in the database, and the AI traverses that structure whenever it is needed. There is no need to push everything into the prompt, and no need to guess by similarity.<SPAN>&nbsp;</SPAN><STRONG>The exact connections live on the graph.</STRONG></P><HR /><H1 id="3-the-problem-enterprise-questions-are-not-single-table" id="toc-hId-902241202">3. The Problem: Enterprise Questions Are Not Single-Table</H1><P class="">A supply chain manager asks a question like this:</P><BLOCKQUOTE dir="auto"><P class=""><EM>"If WorkCenter Z_ASM3 goes down for maintenance, which customers are affected in delivery, and how large is the revenue exposure?"</EM></P></BLOCKQUOTE><P class="">It sounds like a SQL question at first glance. But answering it requires traversing a chain that spans multiple domains:</P><P class=""><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="WorkCenter Product Supply-2026-07-04-060942.svg" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429099i2269D320C391D54C/image-size/large?v=v2&amp;px=999" role="button" title="WorkCenter Product Supply-2026-07-04-060942.svg" alt="WorkCenter Product Supply-2026-07-04-060942.svg" /></span></P><P class="">No single table contains this entire chain. It can be expressed with JOINs, but doing so requires knowing the schema, the join keys, and which tables to include in advance. An AI Agent that receives a natural language question must<SPAN>&nbsp;</SPAN><STRONG>discover</STRONG><SPAN>&nbsp;</SPAN>this path dynamically.</P><H3 id="this-question-is-both-the-starting-point-and-the-destination-of-the-series" id="toc-hId-963893135">This question is both the starting point and the destination of the series</H3><P class="">The question above is not just a demo scenario.<SPAN>&nbsp;</SPAN><STRONG>It is a Competency Question (CQ)</STRONG><SPAN>,</SPAN>&nbsp;the concept Stanford's<SPAN>&nbsp;</SPAN><EM>Ontology Development 101</EM><SPAN>&nbsp;</SPAN>identifies as the starting point of KG design. Before any modeling begins, you write down<SPAN>&nbsp;</SPAN><EM>"What business questions should this KG answer?"</EM></P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span><SPAN>&nbsp;</SPAN><STRONG>The CQ runs through the entire series, from start to finish.</STRONG></P><P class="">A Knowledge Graph cannot be built all at once. The domain is too broad, the data is too large, and the boundaries of<SPAN>&nbsp;</SPAN><EM>"where to stop"</EM><SPAN>&nbsp;</SPAN>are unclear. This is why<SPAN>&nbsp;</SPAN><EM>Ontology Development 101</EM><SPAN>&nbsp;</SPAN>recommends starting from the questions the KG must answer.</P><P class="">The CQ surfaces at three points in this series:</P><OL class=""><LI><STRONG>At the start (Part 2-1, 2-2):</STRONG><SPAN>&nbsp;</SPAN>The CQ determines<SPAN>&nbsp;</SPAN><EM>which classes and relationships are needed.</EM><SPAN>&nbsp;</SPAN>Nouns in the question become classes; verbs become relationships. Rather than modeling every entity, we build only the<SPAN>&nbsp;</SPAN><STRONG>minimum required to answer the question.</STRONG></LI><LI><STRONG>In the middle (Part 3):</STRONG><SPAN>&nbsp;</SPAN>The CQ determines<SPAN>&nbsp;</SPAN><EM>which relationships should be validated first and which data should be prioritized.</EM><SPAN>&nbsp;</SPAN>All relationships matter, but the ones that must be alive first are the chain the CQ points to.</LI><LI><STRONG>At the end (Part 4):</STRONG><SPAN>&nbsp;</SPAN>The CQ becomes the<SPAN>&nbsp;</SPAN><STRONG>quality benchmark for the KG.</STRONG><SPAN>&nbsp;</SPAN>If the CQ produces an answer, the KG is working as intended. If it does not, the gap is visible.<SPAN>&nbsp;</SPAN><STRONG>The CQ effectively becomes the SPARQL test case.</STRONG></LI></OL><P class="">The same question works twice: first as a<SPAN>&nbsp;</SPAN><EM>design specification</EM>, then as a<SPAN>&nbsp;</SPAN><EM>quality benchmark</EM>. This is also why a KG is rarely complete in a single pass: every new CQ triggers the same cycle again.</P></BLOCKQUOTE><P class="">This is why the Z_ASM3 question keeps reappearing throughout the series. Part 2 determines which classes (WorkCenter, FinishedGood, …) and relationships (producesProduct, …) are required to answer it. Part 3 validates those relationships against real data. Part 4 confirms that the Agent actually answers the question through SPARQL and SQL.<SPAN>&nbsp;</SPAN><STRONG>A single question forms the spine of the entire series.</STRONG></P><P class="">And this is precisely the problem Knowledge Graph was designed for: storing a chain that no single table can hold, and traversing it dynamically.</P><HR /><H1 id="4-vector-rag-and-knowledge-graph-different-strengths" id="toc-hId-509214192">4. Vector RAG and Knowledge Graph: Different Strengths</H1><P class="">SAP's official <A href="https://www.sap.com/resources/knowledge-graph" target="_self" rel="noopener noreferrer">Knowledge Graph introduction</A> puts it this way:</P><BLOCKQUOTE dir="auto"><P class=""><EM>"Vector databases help AI find things that are similar; Knowledge graphs help AI understand how things are connected."</EM></P></BLOCKQUOTE><P class="">The two are not competitors. They excel at different things. But when building an AI Agent on top of ERP data, this difference becomes decisive.</P><P class="">Three situations where Vector RAG struggles:</P><P class=""><STRONG>1. Questions that require relationship traversal</STRONG></P><P class=""><EM>"Which finished goods does this part go into?"</EM><SPAN>&nbsp;</SPAN>cannot be answered through similarity search. It requires following a<SPAN>&nbsp;</SPAN><STRONG>structural connection</STRONG>: purchased material → BOM item → finished good. Text embeddings do not represent this connection.</P><P class=""><STRONG>2. Linking data across multiple systems</STRONG></P><P class="">Consider combining S/4HANA procurement data with the World Bank's Logistics Performance Index (LPI) to ask,<SPAN>&nbsp;</SPAN><EM>"Which purchase orders involve suppliers in countries with high logistics risk?"</EM><SPAN>&nbsp;</SPAN>Vector RAG can surface similar passages within each document, but it does not<SPAN>&nbsp;</SPAN><STRONG>explicitly link</STRONG><SPAN>&nbsp;</SPAN>entities from different sources.</P><P class=""><STRONG>3. Precise condition handling</STRONG></P><P class=""><EM>"Suppliers with a risk score below 20"</EM><SPAN>&nbsp;</SPAN>is not a similarity question; it is a filter. Numeric conditions, aggregations, and ordering all require structured queries.</P><P class="">In summary: Vector RAG is strong at document search and semantic similarity. For traversing structural relationships in ERP data, one more layer is required. That layer is the Knowledge Graph.</P><HR /><H1 id="5-what-a-knowledge-graph-solves-relationships-made-explicit" id="toc-hId-312700687">5. What a Knowledge Graph Solves: Relationships Made Explicit</H1><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="WorkCenter Product Supply-2026-07-04-033227.svg" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429095i97A1013C26FAD30C/image-size/large?v=v2&amp;px=999" role="button" title="WorkCenter Product Supply-2026-07-04-033227.svg" alt="WorkCenter Product Supply-2026-07-04-033227.svg" /></span></P><P class="">A Knowledge Graph stores data as<SPAN>&nbsp;</SPAN><STRONG>nodes (entities)</STRONG><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><STRONG>edges (relationships)</STRONG>:</P><PRE><CODE>bdc:WorkCenter_Z_ASM3 bdc:producesProduct bdc:Product_MZ-FG-C900 bdc:WorkCenter_Z_ASM3 bdc:producesProduct bdc:Product_MZ-FG-C950 bdc:WorkCenter_Z_ASM3 bdc:producesProduct bdc:Product_MZ-FG-C990 bdc:WorkCenter_Z_ASM3 bdc:locatedAt bdc:Plant_1010</CODE></PRE><P class="">This is not a table join. It is a structure where<SPAN>&nbsp;</SPAN><STRONG>business-meaningful relationships are stored explicitly.</STRONG></P><P class="">An AI Agent queries this structure with SPARQL:</P><PRE><CODE>SELECT ?product WHERE { bdc:WorkCenter_Z_ASM3 bdc:producesProduct ?product . }</CODE></PRE><P class="">One line. No joins. A query that is understandable without prior schema knowledge.</P><P class="">When numbers are required, the Agent switches to SQL:</P><PRE><CODE><SPAN class="">SELECT</SPAN> "SoldToParty", <SPAN class="">SUM</SPAN>("TotalNetAmount") <SPAN class="">AS</SPAN> revenue <SPAN class="">FROM</SPAN> BDC_DP.SALES_ORDER_ITEM <SPAN class="">WHERE</SPAN> "Material" <SPAN class="">IN</SPAN> (<SPAN class="">'MZ-FG-C900'</SPAN>, <SPAN class="">'MZ-FG-C950'</SPAN>, <SPAN class="">'MZ-FG-C990'</SPAN>) <SPAN class="">GROUP</SPAN> <SPAN class="">BY</SPAN> "SoldToParty";</CODE></PRE><P class="">The SPARQL query identifies three product IDs; SQL aggregates the revenue. The detailed SQL design is covered in a later post in this series. What matters here is the principle itself:<SPAN>&nbsp;</SPAN><STRONG>the two queries divide the work between them.</STRONG></P><P class="">This is the core principle of the Agent we built:<SPAN>&nbsp;</SPAN><STRONG>the KG handles structural relationships, and SQL handles quantitative aggregation.</STRONG></P><HR /><H1 id="6-why-sap-bdc-data-products" id="toc-hId-116187182">6. Why SAP BDC Data Products</H1><P class="">We began Section 1 by noting that Knowledge Graphs generally have a steep learning curve. The good news is that<SPAN>&nbsp;</SPAN><STRONG>SAP data makes KG development significantly more practical</STRONG><SPAN>,</SPAN>&nbsp;thanks to SAP's<SPAN>&nbsp;</SPAN><EM>One Domain Model.</EM></P><P class="">Business entities such as<SPAN>&nbsp;</SPAN><CODE>SalesOrder</CODE>,<SPAN>&nbsp;</SPAN><CODE>Customer</CODE>,<SPAN>&nbsp;</SPAN><CODE>WorkCenter</CODE>, and<SPAN>&nbsp;</SPAN><CODE>Supplier</CODE><SPAN>&nbsp;</SPAN>are already defined in a standardized vocabulary. When these are published as BDC Data Products into HANA Cloud, the CSN metadata is included as well. Half of ontology design, namely <EM>which entities exist and what to call them,</EM>&nbsp;is effectively solved before you start. There is no need to brainstorm entity names on a whiteboard.</P><P class="">The same starting point is available even without going through BDC, when working directly with S/4HANA ERP. CDS View definitions already declare entity names, field annotations, and associations using meaningful business vocabulary. The raw material for an ontology draft is already inside the system. BDC surfaces this as curated Data Products, but teams working directly with S/4HANA can take advantage of the same strength.</P><P class="">LLMs already know this vocabulary. ABAP table names such as<SPAN>&nbsp;</SPAN><CODE>EKKO</CODE><SPAN>&nbsp;</SPAN>or<SPAN>&nbsp;</SPAN><CODE>KNA1</CODE><SPAN>&nbsp;</SPAN>are essentially opaque strings to an LLM, whereas<SPAN>&nbsp;</SPAN><CODE>SalesOrder</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>Customer</CODE><SPAN>&nbsp;</SPAN>appear millions of times in training data. The moment your KG and your LLM share the same vocabulary, writing SPARQL becomes far easier for the Agent.</P><P class="">That said, the original ABAP table names still live inside the SAP system. The question is how SAP BDC Data Products replace that vocabulary with business terminology.</P><H3 id="what-happens-when-you-connect-directly-to-s4hana-tables" id="toc-hId-177839115">What happens when you connect directly to S/4HANA tables</H3><P class="">Consider some core S/4HANA table names:</P><TABLE width="245"><TBODY><TR><TD width="53.9922px"><STRONG>Table</STRONG></TD><TD width="190.008px"><STRONG>Meaning</STRONG></TD></TR><TR><TD width="53.9922px"><CODE>EKKO</CODE></TD><TD width="190.008px">Purchase Order Header</TD></TR><TR><TD width="53.9922px"><CODE>EKPO</CODE></TD><TD width="190.008px">Purchase Order Item</TD></TR><TR><TD width="53.9922px"><CODE>MARA</CODE></TD><TD width="190.008px">Material Master</TD></TR><TR><TD width="53.9922px"><CODE>KNA1</CODE></TD><TD width="190.008px">Customer</TD></TR><TR><TD width="53.9922px"><CODE>CRHD</CODE></TD><TD width="190.008px">Work Center Header</TD></TR></TBODY></TABLE><P class="">These names have lived in SAP systems since the ABAP era. For experienced SAP users, they are immediately recognizable. But<SPAN>&nbsp;</SPAN><STRONG>the names themselves do not convey their meaning intuitively.</STRONG><SPAN>&nbsp;A</SPAN>&nbsp;four- or five-letter abbreviation typically requires separate learning or documentation lookup to map to a business concept.</P><P class="">The same applies to LLMs. An LLM that has encountered these abbreviations frequently in SAP contexts may recognize some of them, but not as reliably as business vocabulary like<SPAN>&nbsp;</SPAN><CODE>SalesOrder</CODE>, which appears millions of times in training data. When building an Agent, a significant share of prompt engineering ends up teaching the LLM that<SPAN>&nbsp;</SPAN><EM>"EKKO is the purchase order header,"</EM>&nbsp;and even then, hallucinations still occur.</P><P class=""><STRONG>SAP BDC Data Products solve this problem at the source.</STRONG></P><H3 id="what-a-bdc-data-product-provides" id="toc-hId--93905759">What a BDC Data Product provides</H3><P class="">SAP Business Data Cloud (BDC) exposes S/4HANA data as curated<SPAN>&nbsp;</SPAN><STRONG>Data Products.</STRONG><SPAN>&nbsp;</SPAN>Each Data Product:</P><UL class=""><LI>Has a<SPAN>&nbsp;</SPAN><STRONG>clear name:</STRONG><SPAN>&nbsp;</SPAN><CODE>SalesOrder</CODE>,<SPAN>&nbsp;</SPAN><CODE>Customer</CODE>,<SPAN>&nbsp;</SPAN><CODE>WorkCenter</CODE>,<SPAN>&nbsp;</SPAN><CODE>Product</CODE></LI><LI>Has a<SPAN>&nbsp;</SPAN><STRONG>CSN (Core Schema Notation) schema:</STRONG><SPAN>&nbsp;</SPAN>including field names, semantic annotations, and FK associations</LI><LI>Is published as a<SPAN>&nbsp;</SPAN><STRONG>HANA Cloud Virtual Table:</STRONG><SPAN>&nbsp;</SPAN>HANA Cloud accesses data in BDC data products through a Virtual Table interface, with no ETL pipeline on your side</LI><LI>Is discoverable through the<SPAN>&nbsp;</SPAN><STRONG>BDC catalog</STRONG></LI></UL><P class="">The actual flow for bringing a Data Product from BDC into HANA Cloud:</P><OL class=""><LI>BDC Catalog &amp; Marketplace → Data Products tab</LI><LI>Select a Data Product → Add Target → Choose HANA Cloud → Share</LI><LI>HANA Cloud Central → Data Product tab → Install</LI></OL><P class="">That is the entire process. There are no complex ETL pipelines. For detailed instructions, please check out this <A href="https://developers.sap.com/tutorials/hana-cloud-data-products-consumption.html" target="_self" rel="noopener noreferrer">SAP tutorial</A>.</P><P class="">After installing the data product, you can also download CSN definitions of it as the screenshot below.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ClaudioJP_0-1783142066073.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429097i0EC1C8A0D9127081/image-size/large?v=v2&amp;px=999" role="button" title="ClaudioJP_0-1783142066073.png" alt="ClaudioJP_0-1783142066073.png" /></span></P><P>The CSN files, downloaded at this stage, will be used in the Part 2-2 of this blog series.</P><H3 id="virtual-table-real-time-access-without-replication" id="toc-hId--290419264">Virtual Table: real-time access without replication</H3><P class=""><SPAN>SAP BDC<STRONG> replicates and curates data from S/4HANA</STRONG> and exposes it as HANA Cloud Virtual Tables. A Virtual Table is not a local copy you manage. BDC handles the data movement from S/4HANA, and HANA Cloud accesses it through the Virtual Table interface. There is no additional ETL pipeline on your side.<BR /></SPAN></P><P class="">For an AI Agent, this has two implications:</P><OL class=""><LI>When generating KG instances, you query the Virtual Table, which reflects BDC's current view of the S/4HANA data.</LI><LI>The Agent's SQL queries run against the same data in data products.</LI></OL><P class="">The trade-off is latency. Cross-system queries are slower than local ones. This is why we place<SPAN>&nbsp;</SPAN><STRONG>stable, smaller master data into the KG</STRONG><SPAN>&nbsp;</SPAN>while handling<SPAN>&nbsp;</SPAN><STRONG>transactional data through Virtual Table SQL</STRONG>, where data freshness matters more than raw speed.</P><HR /><H1 id="7-one-domain-model-an-extensible-vocabulary-foundation" id="toc-hId-99873245">7. One Domain Model: An Extensible Vocabulary Foundation</H1><P class="">BDC Data Products are not simply a name-translation layer. They are built on top of SAP's<SPAN>&nbsp;</SPAN><STRONG>One Domain Model (ODM)</STRONG>. ODM defines business entities in a unified way across SAP products, providing a standard vocabulary at the data model level,<SPAN>&nbsp;</SPAN><CODE>Customer</CODE>,<SPAN>&nbsp;</SPAN><CODE>Product</CODE>,<SPAN>&nbsp;</SPAN><CODE>Supplier</CODE>,<SPAN>&nbsp;</SPAN><CODE>WorkCenter</CODE>. Because the type name itself carries business meaning, the same business concept is identified by the same vocabulary across multiple SAP applications.</P><P class="">The greatest value this standard provides to a KG is<SPAN>&nbsp;</SPAN><STRONG>extensibility.</STRONG></P><P class=""><STRONG>1. When adding other SAP products to the KG.</STRONG><SPAN>&nbsp;</SPAN>Suppose you begin by building a KG with S/4HANA supply chain data. Later, you want to incorporate data from another SAP product into the same KG. If the vocabulary differed across systems, you would need to construct and maintain mapping tables every time. On top of ODM, the starting vocabulary is already shared, so adding a new system<SPAN>&nbsp;</SPAN><STRONG>incrementally</STRONG><SPAN>&nbsp;</SPAN>becomes significantly simpler.</P><P class=""><STRONG>2. When other teams reuse the same KG.</STRONG><SPAN>&nbsp;</SPAN>Because ODM vocabulary is consistent across SAP, KG reuse across teams is straightforward. When another team adds entities from its own domain to an existing KG, the vocabulary and structure do not collide.</P><P class=""><STRONG>3. KG and LLM share the same vocabulary.</STRONG><SPAN>&nbsp;</SPAN>As a side effect, ODM vocabulary (<CODE>Customer</CODE>,<SPAN>&nbsp;</SPAN><CODE>Supplier</CODE>,<SPAN>&nbsp;</SPAN><CODE>WorkCenter</CODE>) is well-represented in LLM training data, so LLMs recognize it reliably. The KG classes we built also use this vocabulary, such as<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter</CODE>,<SPAN>&nbsp;</SPAN><CODE>bdc:Supplier</CODE>. Natural language question → SPARQL → KG result all flow on the same vocabulary.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>One Domain Model turns a KG from "a model you build once and are done with" into "an asset you can continue to extend and connect."</STRONG></P></BLOCKQUOTE><P class=""><span class="lia-unicode-emoji" title=":open_book:">📖</span><SPAN>&nbsp;</SPAN><EM>Reference: SAP Community,<SPAN>&nbsp;</SPAN><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/harnessing-half-a-century-of-knowledge-sap-s-journey-of-enriching-apis-with/ba-p/13578364" target="_blank">"Harnessing Half a Century of Knowledge: SAP's Journey of Enriching APIs with Metadata"</A></EM></P><HR /><H1 id="8-the-division-of-labor-kg-and-sql" id="toc-hId--96640260">8. The Division of Labor: KG and SQL</H1><P class="">The architecture we chose rests on a single principle.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>The KG handles structural relationships. SQL handles quantitative aggregation.</STRONG></P></BLOCKQUOTE><TABLE><TBODY><TR><TD><STRONG>Role</STRONG></TD><TD><STRONG>Technology</STRONG></TD><TD><STRONG>Target&nbsp;Data</STRONG></TD><TD><STRONG><EM>Question Type</EM></STRONG></TD></TR><TR><TD>Relationship traversal</TD><TD>SPARQL over HANA KGE</TD><TD>Master data (WorkCenter, Product, Supplier, Plant)</TD><TD><EM>"What does Z_ASM3 produce?"</EM></TD></TR><TR><TD>Numeric aggregation</TD><TD>SQL over HANA Virtual Table</TD><TD>Transactional data (SalesOrder, PurchaseOrder, BillingDocument)</TD><TD><EM>"How much revenue do those products represent?"</EM></TD></TR></TBODY></TABLE><P class="">SAP BDC Data Products combined with HANA Cloud KGE realize this division of labor on a single database. SPARQL and SQL run on the same HANA instance, and merging their results is handled by the database engine rather than the LLM (we cover this in detail in Part 4 Section 2 with<SPAN>&nbsp;</SPAN><CODE>SPARQL_TABLE</CODE>).</P><P class="">If this principle holds, the next question is<SPAN>&nbsp;</SPAN><STRONG>how to actually build it.</STRONG></P><HR /><H1 id="9-the-full-journey-from-competency-question-to-ai-agent" id="toc-hId--293153765">9. The Full Journey: From Competency Question to AI Agent</H1><P class=""><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="WorkCenter Product Supply-2026-07-04-045016.svg" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429098i20225B1AAAF04133/image-size/large?v=v2&amp;px=999" role="button" title="WorkCenter Product Supply-2026-07-04-045016.svg" alt="WorkCenter Product Supply-2026-07-04-045016.svg" /></span></P><P class=""><STRONG>[0] The Competency Question is both the starting point and the validation benchmark for every step.</STRONG><SPAN>&nbsp;</SPAN>The<SPAN>&nbsp;</SPAN><EM>"CQ runs through the series from start to finish"</EM><SPAN>&nbsp;</SPAN>point emphasized in Section 2 appears directly at positions 0 and 4 of this diagram. The CQ determines<SPAN>&nbsp;</SPAN><EM>which classes and relationships are needed</EM><SPAN>&nbsp;</SPAN>at [2-1] Ontology design, and serves as the test case for<SPAN>&nbsp;</SPAN><EM>whether the KG actually answers the question</EM><SPAN>&nbsp;</SPAN>at [4] Validation. The same question works twice.</P><P class="">KG creation has two parts:</P><UL class=""><LI><STRONG>Ontology (schema):</STRONG><SPAN>&nbsp;</SPAN>What classes and relationships exist. It is created once and rarely changed. Definitions like<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>bdc:producesProduct</CODE><SPAN>&nbsp;</SPAN>belong here.</LI><LI><STRONG>Instance (data):</STRONG><SPAN>&nbsp;</SPAN>Actual WorkCenters, Products, and Suppliers. It is regenerated when data changes. The Virtual Tables are read, and TTL files are produced for the actual entities of the classes declared in the ontology.</LI></UL><P class="">By analogy: the Ontology is the<SPAN>&nbsp;</SPAN><EM>"blueprint,"</EM><SPAN>&nbsp;</SPAN>and the Instance is<SPAN>&nbsp;</SPAN><EM>"the actual building constructed from that blueprint."</EM></P><TABLE><TBODY><TR><TD width="137.688px" height="30px"><STRONG>Step</STRONG></TD><TD width="153.828px" height="30px"><STRONG>Key&nbsp;Question</STRONG></TD><TD width="263.406px" height="30px"><STRONG>HANA KGE Feature </STRONG></TD><TD width="190.578px" height="30px"><STRONG>Covered in</STRONG></TD></TR><TR><TD width="137.688px" height="112px"><STRONG>[0] Competency Question</STRONG></TD><TD width="153.828px" height="112px">What business questions should this KG answer?</TD><TD width="263.406px" height="112px">(design starting point)</TD><TD width="190.578px" height="112px"><STRONG>Part 2-1 / 2-2 (starting point) + Part 4 (validation benchmark)</STRONG></TD></TR><TR><TD width="137.688px" height="85px"><STRONG>[1] Data Integration</STRONG></TD><TD width="153.828px" height="85px">How do I bring a BDC Data Product into HANA?</TD><TD width="263.406px" height="85px">Virtual Table (no ETL)</TD><TD width="190.578px" height="85px">Part 1 (this post)</TD></TR><TR><TD width="137.688px" height="85px"><STRONG>[2-1] KG Creation — Ontology</STRONG></TD><TD width="153.828px" height="85px">How do I design classes and relationships?</TD><TD width="263.406px" height="85px">(design phase)</TD><TD width="190.578px" height="85px">Part 2-1 / 2-2</TD></TR><TR><TD width="137.688px" height="85px"><STRONG>[2-2] KG Creation — Instance</STRONG></TD><TD width="153.828px" height="85px">How do I generate TTL from real data?</TD><TD width="263.406px" height="85px">(generation phase)</TD><TD width="190.578px" height="85px">Part 3</TD></TR><TR><TD width="137.688px" height="85px"><STRONG>[3] Load</STRONG></TD><TD width="153.828px" height="85px">How do I import TTL into HANA KGE?</TD><TD width="263.406px" height="85px"><STRONG>Triple Store</STRONG>: a dedicated graph store separated from relational data</TD><TD width="190.578px" height="85px">Part 3</TD></TR><TR><TD width="137.688px" height="85px"><STRONG>[4] Validation</STRONG></TD><TD width="153.828px" height="85px">How do I verify that the loaded KG is correct?</TD><TD width="263.406px" height="85px"><STRONG>SHACL</STRONG>: W3C standard for graph data validation</TD><TD width="190.578px" height="85px">Part 3 (introduced briefly)</TD></TR><TR><TD width="137.688px" height="112px"><STRONG>[5] AI Agent Development</STRONG></TD><TD width="153.828px" height="112px">How do I combine SPARQL + SQL?</TD><TD width="263.406px" height="112px"><STRONG>SPARQL_TABLE</STRONG>: run SPARQL directly inside SQL.<SPAN>&nbsp;</SPAN><STRONG>Inference</STRONG>: derive relationships from OWL rules without explicit triples</TD><TD width="190.578px" height="112px">Part 4 (SPARQL_TABLE in depth / Inference introduced briefly)</TD></TR></TBODY></TABLE><BLOCKQUOTE dir="auto"><P class=""><STRONG>[0] CQ is not a step. It is the baseline for every step.</STRONG><SPAN>&nbsp;</SPAN>It is not written once and set aside. At ontology design, it answers<SPAN>&nbsp;</SPAN><EM>which classes are required;</EM><SPAN>&nbsp;</SPAN>at validation, it tests<SPAN>&nbsp;</SPAN><EM>whether the KG actually answers the question.</EM><SPAN>&nbsp;</SPAN>The same question runs from the beginning of the series to the end. The Section 2 point, <EM>"the CQ runs through the series from start to finish,"</EM>&nbsp;is also why the [0] row in this table appears in both Part 2-1/2-2 and Part 4.</P></BLOCKQUOTE><P class="">A one-line preview of each feature:</P><UL class=""><LI><STRONG>Triple Store</STRONG>: Stores KG data in a space completely separate from relational tables, managed at graph granularity, with independent permission control.</LI><LI><STRONG>SHACL</STRONG><SPAN>&nbsp;</SPAN><EM>(not covered in detail in this series)</EM>:<SPAN>&nbsp;</SPAN><EM>"This field must exist," "This value must be a number."</EM><SPAN>&nbsp;</SPAN>Define the rules once, and they are checked automatically as data is loaded. AI answer quality follows from data quality.</LI><LI><STRONG>SPARQL_TABLE</STRONG>: Query the KG with one line in the SQL Console, for example <CODE>SELECT * FROM SPARQL_TABLE('...')</CODE>. No separate SPARQL endpoint is required.</LI><LI><STRONG>Inference</STRONG><SPAN>&nbsp;</SPAN><EM>(not covered in detail in this series)</EM>: Derive relationships through RDFS/OWL rules without explicit triples. Declare<SPAN>&nbsp;</SPAN><CODE>Manager = Leader</CODE><SPAN>&nbsp;</SPAN>in the ontology, and Alice (a Manager) is automatically queryable as a Leader without modifying the data.</LI></UL><P class="">These four features form the core of HANA Cloud KGE. How each one is applied in practice is covered in the corresponding post in this series.</P><HR /><H2 id="whats-next" id="toc-hId--783070277">What's Next</H2><P class="">If this principle holds, the next question is how to actually build it.</P><P class=""><STRONG>Part 2-1 – Ontology Standards and Concepts</STRONG><SPAN>&nbsp;</SPAN>covers the standards and concepts: what an ontology is, Class vs Instance, and how the W3C standards (Direct Mapping and R2RML) define the translation from relational data to RDF.<SPAN>&nbsp;</SPAN><STRONG>Part 2-2 – Designing the Ontology for SAP Data</STRONG><SPAN>&nbsp;</SPAN>then walks through the practical design decisions we made for SAP data, including what to place in the KG versus leave in SQL, Class Partition, the resulting ontology, and how we used AI to draft it. The two posts work as a pair, but Part 2-2 is also designed to stand on its own.</P><HR /><P class=""><EM>Stack: SAP BDC Data Products · HANA Cloud KGE · FastAPI · React · Claude (Anthropic / AI Core)</EM><SPAN>&nbsp;</SPAN><EM>Code:<SPAN>&nbsp;</SPAN><A href="https://github.com/claudiopark86/sap-bdc-kge-agent-workshop" target="_blank" rel="noopener nofollow noreferrer">github.com/claudiopark86/sap-bdc-kge-agent-workshop</A></EM></P><HR /><H2 id="references" id="toc-hId--979583782">References</H2><UL class=""><LI><STRONG>Ontology Development 101: A Guide to Creating Your First Ontology</STRONG><SPAN>&nbsp;</SPAN>— Noy &amp; McGuinness, Stanford 2001. The standard introduction to ontology design methodology. The source for concepts covered throughout this series, including Competency Questions, Class vs Instance, and iterative design.<SPAN>&nbsp;</SPAN><A href="https://protege.stanford.edu/publications/ontology_development/ontology101.pdf" target="_blank" rel="noopener nofollow noreferrer">https://protege.stanford.edu/publications/ontology_development/ontology101.pdf</A></LI><LI><STRONG>Harnessing Half a Century of Knowledge: SAP's Journey of Enriching APIs with Metadata</STRONG><SPAN>&nbsp;</SPAN>— SAP Community Blog. SAP's evolved metadata strategy and ODM vision.<SPAN>&nbsp;</SPAN><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/harnessing-half-a-century-of-knowledge-sap-s-journey-of-enriching-apis-with/ba-p/13578364" target="_blank">https://community.sap.com/t5/technology-blog-posts-by-sap/harnessing-half-a-century-of-knowledge-sap-s-journey-of-enriching-apis-with/ba-p/13578364</A></LI></UL> 2026-07-04T15:44:34.848000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-2-1-ontology-standards/ba-p/14417385 Knowledge Graph Agent on SAP HANA Cloud Series – Part 2-1: Ontology Standards and Concepts 2026-07-04T15:45:03.995000+02:00 ClaudioJP https://community.sap.com/t5/user/viewprofilepage/user-id/1509109 <P class="">This blog is part of a blog series on building AI Agents with SAP BDC (Business Data Cloud) Data Products and SAP HANA Cloud Knowledge Graph Engine:</P><UL><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-1-why-knowledge-graph/ba-p/14417381" target="_blank">Knowledge Graph Agent on SAP HANA Cloud Series – Part 1: Why Knowledge Graph for SAP Data</A></LI><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-2-1-ontology-standards/ba-p/14417385" target="_blank">Knowledge Graph Agent on SAP HANA Cloud Series – Part 2-1: Ontology Standards and Concepts</A></LI><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-2-2-designing-the/ba-p/14274465" target="_blank">Knowledge Graph Agent on SAP HANA Cloud Series – Part 2-2: Designing the Ontology for SAP Data</A>&nbsp;</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series – Part 3: Data Pipeline from BDC to HANA Cloud KGE</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series – Part 4: Building an Agent with SPARQL and SQL</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series&nbsp;– Part 5: Lessons Learned</LI></UL><P><SPAN>Enterprise systems are vast, with thousands of tables, hundreds of business processes, and countless APIs, all interconnected in ways only domain experts fully grasp. For AI Agents to be genuinely useful in this environment, they need to understand not just what the data is, but how it all connects: which Product belongs to which Sales Order, which WorkCenter is located in which Plant, which Supplier serves which material. This understanding does not come from the data itself. It needs a dedicated semantic layer that makes business meaning explicit and machine-readable. Knowledge Graphs provide exactly that layer. SAP HANA Cloud now offers a native Knowledge Graph Engine to build it on top of SAP BDC Data Products. This series walks through the full journey: from why this layer matters, to how to design and build it, to how an AI Agent uses it to answer real business questions.</SPAN></P><P class="">This post covers the concepts and standards of ontology design. It explains what an ontology is, why it works well with LLMs, what the seven-step methodology from Stanford's<SPAN>&nbsp;</SPAN><EM>Ontology Development 101</EM><SPAN>&nbsp;</SPAN>proposes, and what the W3C standards (Direct Mapping and R2RML) define when translating relational data into RDF. This is the most theoretically detailed post in the series, providing the vocabulary and references that underpin the design decisions in later posts.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span><SPAN>&nbsp;</SPAN><STRONG>Reading guide.</STRONG><SPAN>&nbsp;</SPAN>This post goes deep into theory and W3C standards. If you would rather see the practical SAP BDC decisions right away, you can skip ahead to<SPAN>&nbsp;</SPAN><STRONG>Part 2-2 – Designing the Ontology for SAP Data</STRONG>. Part 2-2 is structured so that you can follow it without reading this post to the end, and you can return here when you want the background.</P></BLOCKQUOTE><P class=""><STRONG>Table of contents:</STRONG></P><OL class=""><LI>What is an Ontology<UL class=""><LI>1.1 Core concepts</LI><LI>1.2 Class vs Instance</LI><LI>1.3 Property: connecting nodes and carrying values (ObjectProperty / DatatypeProperty / Domain &amp; Range)</LI><LI>1.4 Design principles (IRI, Competency Questions)</LI><LI>1.5 Labels, multilingual text, and synonyms (rdfs:label, language tags, skos:altLabel)</LI></UL></LI><LI>The Seven Steps of Ontology 101: the process we followed</LI><LI>From RDB to RDF: what the W3C standards define<UL class=""><LI>3.1 Class definition</LI><LI>3.2 Instance data identification</LI><LI>3.3 Data attribute representation</LI><LI>3.4 Relationship representation</LI><LI>Combined example</LI></UL></LI><LI>Three domain decisions we made on top of the standards</LI></OL><HR /><H2 id="what-this-post-covers" id="toc-hId-1817377945">What this post covers</H2><P class="">This post addresses the<SPAN>&nbsp;</SPAN><STRONG>theory and standards</STRONG><SPAN>&nbsp;</SPAN>within the series. It introduces what an ontology is, the seven-step methodology from Stanford's<SPAN>&nbsp;</SPAN><EM>Ontology Development 101</EM>, and the two W3C standards (Direct Mapping and R2RML) that define how relational data is mapped to RDF, including what they automate and where domain decisions are required.</P><P class="">This material is the<SPAN>&nbsp;</SPAN><STRONG>background knowledge</STRONG><SPAN>&nbsp;</SPAN>for the decisions that follow (Class Partition, KG vs SQL classification, Relationship SQL Views). The decisions themselves are covered from Part 2-2 onward, so if you want to move directly to practical decisions, you can skip ahead to Part 2-2.</P><P class=""><STRONG>One practical use:</STRONG><SPAN>&nbsp;</SPAN>the content of this post can also be provided to an LLM as a prompt. The seven-step methodology from Ontology 101, the W3C mapping rules (Table → Class, Row → Instance, and so on), and the conventions for defining classes and properties using RDFS/OWL vocabulary. Feeding this standards guide together with your domain information to a generative model produces ontology drafts that are much closer to W3C standards. This is what determines the quality of the<SPAN>&nbsp;</SPAN><EM>"drafting with AI"</EM><SPAN>&nbsp;</SPAN>approach we cover in Part 2-2.</P><HR /><H1 id="1-what-is-an-ontology" id="toc-hId-1491781721">1. What is an Ontology</H1><H3 id="11-core-concepts" id="toc-hId-1553433654">1.1 Core concepts</H3><P class="">Before building a Knowledge Graph, it is important to understand what an ontology is. Stated simply:</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>An ontology is an explicit definition of "what kinds of things exist in this world, and how they are connected."</STRONG></P></BLOCKQUOTE><P class="">For an ontology in the supply chain domain:</P><UL class=""><LI>These kinds of things exist: WorkCenter, Product, Supplier, Plant</LI><LI>They are connected like this: WorkCenters produce Products, Suppliers provide Products</LI></UL><P class="">An ontology file is this information written in a machine-readable format (RDF/OWL).</P><P class="">Why build one?<SPAN>&nbsp;</SPAN><EM>Ontology Development 101</EM><SPAN>&nbsp;</SPAN>puts it this way:</P><BLOCKQUOTE dir="auto"><P class=""><EM>"Developing an ontology is akin to defining a set of data and their structure for other programs to use."</EM></P></BLOCKQUOTE><P class="">The goal is not the ontology itself. The goal is<SPAN>&nbsp;</SPAN><STRONG>to define the structure and vocabulary of data that an AI Agent or application will consume.</STRONG><SPAN>&nbsp;</SPAN>At its core, the value is<SPAN>&nbsp;</SPAN><STRONG>common understanding:</STRONG>&nbsp;making it possible for humans, LLMs, and software agents to interpret the same concepts through the same vocabulary.</P><P class="">This is where SAP BDC and One Domain Model come in. ODM already defines a<SPAN>&nbsp;</SPAN><STRONG>standard vocabulary</STRONG><SPAN>&nbsp;across S/4HANA, including</SPAN><SPAN>&nbsp;</SPAN><CODE>WorkCenter</CODE>,<SPAN>&nbsp;</SPAN><CODE>SalesOrder</CODE>, and<SPAN>&nbsp;</SPAN><CODE>Supplier</CODE>. When you build an ontology, you use this vocabulary directly. The LLM already knows what these words mean. The KG is structured with the same vocabulary. The Agent writes SPARQL using the same vocabulary.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>The "common understanding" that Ontology 101 advocates is, in the case of SAP data, already provided by One Domain Model. Half of ontology design is done before you begin.</STRONG></P></BLOCKQUOTE><PRE><CODE># Ontology example (Turtle format) bdc:WorkCenter a owl:Class ; rdfs:label "WorkCenter" ; rdfs:comment "Production work center" . bdc:Product a owl:Class ; rdfs:label "Product" ; rdfs:comment "Material/product" . bdc:producesProduct a owl:ObjectProperty ; rdfs:domain bdc:WorkCenter ; rdfs:range bdc:Product ; rdfs:label "produces product" .</CODE></PRE><H3 id="12-class-vs-instance-the-first-distinction" id="toc-hId-1356920149">1.2 Class vs Instance: the first distinction</H3><P class="">When first encountering ontologies, the difference between<SPAN>&nbsp;</SPAN><STRONG>Class</STRONG><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><STRONG>Instance</STRONG><SPAN>&nbsp;</SPAN>is often confusing. This distinction is the starting point for the entire KG structure.</P><BLOCKQUOTE dir="auto"><UL class=""><LI><STRONG>Class</STRONG><SPAN>&nbsp;</SPAN>= concept.<SPAN>&nbsp;</SPAN><EM>"There is a kind of thing called WorkCenter."</EM></LI><LI><STRONG>Instance</STRONG><SPAN>&nbsp;</SPAN>= actual data.<SPAN>&nbsp;</SPAN><EM>"Z_ASM3 is an actual WorkCenter that exists."</EM></LI></UL></BLOCKQUOTE><P class="">By analogy with a BDC Virtual Table:</P><TABLE><TBODY><TR><TD><STRONG>KG concept</STRONG></TD><TD><STRONG>&nbsp;Example</STRONG></TD><TD><STRONG>Virtual Table analogy</STRONG></TD></TR><TR><TD><STRONG>Class</STRONG></TD><TD><CODE>bdc:WorkCenter</CODE></TD><TD>Table schema (column definitions)</TD></TR><TR><TD><STRONG>Instance</STRONG></TD><TD><CODE>bdc:WorkCenter_Z_ASM3</CODE></TD><TD>One row in the table</TD></TR></TBODY></TABLE><P class="">This is why there are two files:</P><UL class=""><LI><CODE>bdc_dp_ontology.ttl</CODE>: class and relationship definitions (schema)</LI><LI><CODE>bdc_dp_instances.ttl</CODE>: actual WorkCenter, Product, and Supplier data</LI></UL><P class="">There are as many instances as there are rows in the Virtual Table. With the demo data containing 129 WorkCenters, there are 129 instances.</P><H3 id="13-property-connecting-nodes-and-carrying-values" id="toc-hId-1160406644">1.3 Property: connecting nodes and carrying values</H3><P class="">If Class and Instance describe<SPAN>&nbsp;</SPAN><EM>"what exists,"</EM><SPAN>&nbsp;</SPAN>Property describes<SPAN>&nbsp;</SPAN><EM>"how things are connected and what values they hold."</EM><SPAN>&nbsp;</SPAN>RDF/OWL distinguishes two types of property:</P><TABLE><TBODY><TR><TD><STRONG>Type</STRONG></TD><TD><STRONG>Value form&nbsp;</STRONG></TD><TD><STRONG>Example</STRONG></TD></TR><TR><TD><STRONG><CODE>owl:ObjectProperty</CODE></STRONG><SPAN>&nbsp;</SPAN>(relationship)</TD><TD>IRI of another node</TD><TD><CODE>bdc:WorkCenter_Z_ASM3 bdc:producesProduct bdc:Product_MZ-FG-C900</CODE></TD></TR><TR><TD><STRONG><CODE>owl:DatatypeProperty</CODE></STRONG><SPAN>&nbsp;</SPAN>(attribute)</TD><TD>literal (string / number / date)</TD><TD><CODE>bdc:WorkCenter_Z_ASM3 bdc:workCenterName "Z_ASM3"</CODE></TD></TR></TBODY></TABLE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span><SPAN>&nbsp;</SPAN><STRONG>If this is confusing:</STRONG><SPAN>&nbsp;</SPAN>think of<SPAN>&nbsp;</SPAN><STRONG>DatatypeProperty as "a regular column that does not lead to another table."</STRONG><SPAN>&nbsp;</SPAN>A relational column such as<SPAN>&nbsp;</SPAN><CODE>WORK_CENTER.WorkCenterName</CODE>, which simply holds a row's attribute, becomes a DatatypeProperty like<SPAN>&nbsp;</SPAN><CODE>bdc:workCenterName</CODE><SPAN>&nbsp;</SPAN>in the KG. Conversely,<SPAN>&nbsp;</SPAN><STRONG>ObjectProperty resembles "a foreign key that points to another table."</STRONG><SPAN>&nbsp;</SPAN>A column such as<SPAN>&nbsp;</SPAN><CODE>WORK_CENTER.PlantID</CODE>, which points to another row, becomes an ObjectProperty like<SPAN>&nbsp;</SPAN><CODE>bdc:locatedAt</CODE><SPAN>&nbsp;</SPAN>in the KG.</P></BLOCKQUOTE><P class="">This distinction is<SPAN>&nbsp;</SPAN><STRONG>what makes a graph a graph.</STRONG><SPAN>&nbsp;</SPAN>With only DatatypeProperties, each node would carry its own attributes but would not link to other nodes. They would be closer to isolated records than a connected graph. The moment ObjectProperty connects a node to another node, the records become a graph, and SPARQL can traverse chains such as<SPAN>&nbsp;</SPAN><EM>"WorkCenter → Product → Customer."</EM></P><P class="">In Turtle, both are declared with the same<SPAN>&nbsp;</SPAN><CODE>rdf:type</CODE><SPAN>&nbsp;</SPAN>syntax:</P><PRE><CODE>bdc:producesProduct a owl:ObjectProperty ; rdfs:domain bdc:WorkCenter ; rdfs:range bdc:FinishedGood . bdc:workCenterName a owl:DatatypeProperty ; rdfs:domain bdc:WorkCenter ; rdfs:range xsd:string .</CODE></PRE><H4 id="domain--range-applying-type-constraints-to-properties" id="toc-hId-1092975858">Domain &amp; Range: applying type constraints to properties</H4><P class="">Each property can carry type constraints through<SPAN>&nbsp;</SPAN><CODE>rdfs:domain</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>rdfs:range</CODE>:</P><UL class=""><LI><STRONG><CODE>rdfs:domain</CODE></STRONG>: the class expected to appear as the<SPAN>&nbsp;</SPAN><STRONG>subject</STRONG><SPAN>&nbsp;</SPAN>of this property.<SPAN>&nbsp;</SPAN><EM>"<CODE>producesProduct</CODE><SPAN>&nbsp;</SPAN>is intended to be used with<SPAN>&nbsp;</SPAN><CODE>WorkCenter</CODE><SPAN>&nbsp;</SPAN>as its subject."</EM></LI><LI><STRONG><CODE>rdfs:range</CODE></STRONG>:<UL class=""><LI>For ObjectProperty: the class that can serve as the<SPAN>&nbsp;</SPAN><STRONG>object</STRONG>.<SPAN>&nbsp;</SPAN><EM>"The value of<SPAN>&nbsp;</SPAN><CODE>producesProduct</CODE><SPAN>&nbsp;</SPAN>is a<SPAN>&nbsp;</SPAN><CODE>FinishedGood</CODE>."</EM></LI><LI>For DatatypeProperty: the<SPAN>&nbsp;</SPAN><STRONG>XSD type</STRONG><SPAN>&nbsp;</SPAN>of the value.<SPAN>&nbsp;</SPAN><CODE>xsd:string</CODE>,<SPAN>&nbsp;</SPAN><CODE>xsd:integer</CODE>,<SPAN>&nbsp;</SPAN><CODE>xsd:date</CODE>, and so on.</LI></UL></LI></UL><P class="">Without these modeling guardrails, SPARQL queries may return invalid results, for example, a<SPAN>&nbsp;</SPAN><CODE>Supplier</CODE><SPAN>&nbsp;</SPAN>producing a product, or a<SPAN>&nbsp;</SPAN><CODE>Plant</CODE><SPAN>&nbsp;</SPAN>being produced. Domain and Range help both humans and LLMs use the intended classes when writing SPARQL queries.</P><P class="">Practical guidance (Ontology 101):</P><UL class=""><LI>Use<SPAN>&nbsp;</SPAN><STRONG>the most general class</STRONG><SPAN>&nbsp;</SPAN>for Domain and Range. If you want to include every Product, use<SPAN>&nbsp;</SPAN><CODE>bdc:Product</CODE><SPAN>&nbsp;</SPAN>instead of<SPAN>&nbsp;</SPAN><CODE>bdc:FinishedGood</CODE>.</LI><LI>Do not make the range too broad either (setting it to<SPAN>&nbsp;</SPAN><CODE>owl:Thing</CODE><SPAN>&nbsp;</SPAN>carries no meaning).</LI><LI>All subclasses are automatically included. If the range is<SPAN>&nbsp;</SPAN><CODE>bdc:Product</CODE>, its subclasses (FinishedGood/SemiFinished/RawMaterial) are included as well.</LI></UL><H4 id="edge-case-datatypeproperty-vs-objectproperty" id="toc-hId-896462353">Edge case: DatatypeProperty vs ObjectProperty</H4><P class="">There are cases where it is not obvious whether a value should be a<SPAN>&nbsp;</SPAN><EM>"plain string"</EM><SPAN>&nbsp;</SPAN>or<SPAN>&nbsp;</SPAN><EM>"a separate node."</EM><SPAN>&nbsp;</SPAN>For example, the<SPAN>&nbsp;</SPAN><STRONG>country</STRONG><SPAN>&nbsp;</SPAN>of a supplier:</P><UL class=""><LI>If a simple value is sufficient →<SPAN>&nbsp;</SPAN><CODE>owl:DatatypeProperty</CODE><SPAN>&nbsp;</SPAN>(<CODE>bdc:country "DE"^^xsd:string</CODE>)</LI><LI>If multiple systems or multiple classes reference the same country, or if the country itself needs a label or description → model the country as a separate node and use<SPAN>&nbsp;</SPAN><CODE>owl:ObjectProperty</CODE><SPAN>&nbsp;</SPAN>(<CODE>bdc:locatedInCountry bdc:Country_DE</CODE>)</LI></UL><P class="">The latter is the pattern that turns a KG into an<SPAN>&nbsp;</SPAN><EM>"extensible asset,"</EM><SPAN>&nbsp;</SPAN>but for simple representations, a DatatypeProperty is sufficient. This is a<SPAN>&nbsp;</SPAN><STRONG>domain decision</STRONG>.</P><H3 id="14-design-principles-iri-competency-questions" id="toc-hId-570866129">1.4 Design principles: IRI, Competency Questions</H3><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span><SPAN>&nbsp;</SPAN><STRONG>IRI notation.</STRONG><SPAN>&nbsp;</SPAN>In RDF, every resource (class, property, instance) is identified by an<SPAN>&nbsp;</SPAN><STRONG>IRI</STRONG><SPAN>&nbsp;</SPAN>(Internationalized Resource Identifier). You can think of it as a URL. The two are nearly the same for this discussion. Originally, it appears in the full URL form, such as<SPAN>&nbsp;</SPAN><CODE>&lt;<A href="http://sap.com/bdc/WorkCenter" target="_blank" rel="noopener noreferrer">http://sap.com/bdc/WorkCenter</A>&gt;</CODE>.</P><P class=""><STRONG>Why include the domain (URL)?</STRONG><SPAN>&nbsp;</SPAN>To prevent collisions when the same name is used with different meanings across domains. For example, the<SPAN>&nbsp;</SPAN><CODE>Customer</CODE><SPAN>&nbsp;</SPAN>in one system and the<SPAN>&nbsp;</SPAN><CODE>Customer</CODE><SPAN>&nbsp;</SPAN>in another system may represent different concepts; embedding the domain in the IRI makes the two resources distinguishable. Both can be loaded into the same KG without collision, and when two KGs are later integrated, you can decide explicitly how to map them.</P><P class="">Writing the full URL every time becomes verbose, so Turtle allows<SPAN>&nbsp;</SPAN><STRONG>prefix declarations</STRONG><SPAN>&nbsp;</SPAN>as a shorthand. This series uses<SPAN>&nbsp;</SPAN><CODE>bdc:</CODE><SPAN>&nbsp;</SPAN>as the single prefix for the SAP BDC domain.</P><PRE><CODE>@prefix bdc: &lt;http://sap.com/bdc/&gt; . @prefix rdfs: &lt;http://www.w3.org/2000/01/rdf-schema#&gt; . @prefix owl: &lt;http://www.w3.org/2002/07/owl#&gt; . @prefix xsd: &lt;http://www.w3.org/2001/XMLSchema#&gt; .</CODE></PRE><P class="">From here on,<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter</CODE><SPAN>&nbsp;</SPAN>(=<SPAN>&nbsp;</SPAN><CODE>&lt;<A href="http://sap.com/bdc/WorkCenter" target="_blank" rel="noopener noreferrer">http://sap.com/bdc/WorkCenter</A>&gt;</CODE>) and<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter_Z_ASM3</CODE><SPAN>&nbsp;</SPAN>(=<SPAN>&nbsp;</SPAN><CODE>&lt;<A href="http://sap.com/bdc/WorkCenter_Z_ASM3" target="_blank" rel="noopener noreferrer">http://sap.com/bdc/WorkCenter_Z_ASM3</A>&gt;</CODE>) can be written in this short form. Classes, relationships, properties, and instances are distinguished by<SPAN>&nbsp;</SPAN><STRONG>naming conventions</STRONG>:</P><BR /><TABLE><TBODY><TR><TD><STRONG>Type</STRONG></TD><TD><STRONG>Convention</STRONG></TD><TD><STRONG>Example</STRONG></TD></TR><TR><TD>Class</TD><TD>PascalCase</TD><TD><CODE>bdc:WorkCenter</CODE>,<SPAN>&nbsp;</SPAN><CODE>bdc:Customer</CODE></TD></TR><TR><TD>Relationship (ObjectProperty)</TD><TD>camelCase, verb phrase</TD><TD><CODE>bdc:producesProduct</CODE>,<SPAN>&nbsp;</SPAN><CODE>bdc:hasAddress</CODE></TD></TR><TR><TD>Attribute (DatatypeProperty)</TD><TD>camelCase</TD><TD><CODE>bdc:country</CODE>,<SPAN>&nbsp;</SPAN><CODE>bdc:customerID</CODE></TD></TR><TR><TD>Instance</TD><TD><CODE>ClassName_identifier</CODE></TD><TD><CODE>bdc:WorkCenter_Z_ASM3</CODE>,<SPAN>&nbsp;</SPAN><CODE>bdc:Customer_0001000010</CODE></TD></TR></TBODY></TABLE><P class="">This convention follows the practice of community-standard RDF vocabularies such as schema.org and FOAF. Most code examples in this series use the abbreviated form. The W3C standard examples in Section 3 (<CODE>&lt;People/ID=7&gt;</CODE>) use full IRIs without prefixes. Both forms represent the same IRI notation.</P><P class=""><STRONG>For extension: separate prefixes by domain.</STRONG><SPAN>&nbsp;</SPAN>We start with the single SAP BDC domain, so a single<SPAN>&nbsp;</SPAN><CODE>bdc:</CODE><SPAN>&nbsp;</SPAN>prefix is enough. As the KG expands into other domains, the standard RDF convention is to introduce a separate prefix per domain. For example, when integrating data from another SAP product, or external systems such as logistics information or third-party master data, those would be placed under a different prefix and domain URL so the two vocabularies do not collide. Rather than packing everything into<SPAN>&nbsp;</SPAN><CODE>bdc:</CODE>, prefixes are added along domain boundaries. This is the prefix-level realization of the<SPAN>&nbsp;</SPAN><EM>"KG as an extensible asset"</EM><SPAN>&nbsp;</SPAN>point emphasized in Part 1 Section 6.</P></BLOCKQUOTE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":open_book:">📖</span><SPAN>&nbsp;</SPAN><STRONG>Terminology note:</STRONG><SPAN>&nbsp;</SPAN>In Description Logic and KG literature, the two layers are called<SPAN>&nbsp;</SPAN><STRONG>T-Box</STRONG><SPAN>&nbsp;</SPAN>(<EM>Terminological Box</EM>, definitions of classes and relationships, the schema) and<SPAN>&nbsp;</SPAN><STRONG>A-Box</STRONG><SPAN>&nbsp;</SPAN>(<EM>Assertion Box</EM>, the actual instance facts). The two TTL files in our project correspond exactly to these layers:<SPAN>&nbsp;</SPAN><CODE>bdc_dp_ontology.ttl</CODE><SPAN>&nbsp;</SPAN>is the T-Box, and<SPAN>&nbsp;</SPAN><CODE>bdc_dp_instances.ttl</CODE><SPAN>&nbsp;</SPAN>is the A-Box.</P></BLOCKQUOTE><H4 id="competency-questions-write-the-questions-before-designing" id="toc-hId-503435343">Competency Questions: write the questions before designing</H4><P class="">Ontology Development 101 (Noy &amp; McGuinness, Stanford 2001) emphasizes one core principle:</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>"Before designing an ontology, write down the questions this KG must answer."</STRONG></P></BLOCKQUOTE><P class="">These are called<SPAN>&nbsp;</SPAN><STRONG>Competency Questions</STRONG>. Defining the questions first determines naturally which classes and relationships are needed.</P><P class="">The Competency Questions for our project:</P><TABLE><TBODY><TR><TD><STRONG>Question</STRONG></TD><TD><STRONG>What is needed</STRONG></TD></TR><TR><TD><EM>"What finished goods does Z_ASM3 produce?"</EM></TD><TD><P><CODE>bdc:WorkCenter</CODE></P><P><CODE>bdc:producesProduct</CODE></P><P><CODE>bdc:FinishedGood</CODE></P></TD></TR><TR><TD><EM>"What is the per-customer revenue exposure for these finished goods?"</EM></TD><TD>SQL (No KG required. Transactional data)</TD></TR><TR><TD><EM>"What products does this supplier provide?"</EM></TD><TD><P><CODE>bdc:Supplier</CODE></P><P><CODE>bdc:fromSupplier</CODE></P><P><SPAN>&nbsp;</SPAN><CODE>bdc:forProduct</CODE></P><P><SPAN>&nbsp;</SPAN><CODE>bdc:PurchasingSourceList</CODE></P></TD></TR><TR><TD><EM>"Which finished goods does this raw material go into?"</EM></TD><TD><CODE>bdc:hasComponent</CODE><SPAN>&nbsp;</SPAN>(Product → Product)</TD></TR></TBODY></TABLE><P class="">Once the questions are written, three things are determined:</P><OL class=""><LI><STRONG>Which classes are needed</STRONG>: the nouns in the question.</LI><LI><STRONG>Which relationships are needed</STRONG>: the verbs in the question.</LI><LI><STRONG>What goes into the KG versus what stays in SQL</STRONG>: relationships requiring SPARQL traversal go into the KG; numeric aggregation stays in SQL.</LI></OL><BLOCKQUOTE dir="auto"><P class="">These questions later serve as SPARQL test cases. After loading, if SPARQL returns the correct answer for each question, the ontology is well designed.</P></BLOCKQUOTE><H3 id="15-labels-multilingual-text-and-synonyms-rdfslabel-language-tags-skosaltlabel" id="toc-hId-177839119">1.5 Labels, multilingual text, and synonyms: rdfs:label, language tags, skos:altLabel</H3><P><SPAN><span class="lia-unicode-emoji" title=":light_bulb:">💡</span>&nbsp;</SPAN><STRONG>Reading note.</STRONG><SPAN>&nbsp;This section is longer because labels, comments, and synonyms directly affect how well an AI Agent interprets KG results. If you already know RDFS/SKOS annotations, you can skim this section and continue with Section 2.</SPAN></P><H4 id="rdfslabel-and-rdfscomment-the-basics" id="toc-hId--387308762">rdfs:label and rdfs:comment: the basics</H4><P class="">In a KG, attaching<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>to every Class, Property, and Instance is standard practice. Without labels, SPARQL results return only URIs such as<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter_Z_ASM3</CODE>, which are hard for an LLM to interpret.<SPAN>&nbsp;</SPAN><CODE>rdfs:comment</CODE><SPAN>&nbsp;</SPAN>adds a short description that helps the LLM understand the role of each class or property when ontology context is provided. Practical authoring rules are covered in Part 2-2.</P><H4 id="multilingual-language-tags" id="toc-hId--583822267">Multilingual: language tags</H4><P class="">Attaching a language tag to the same property supports multiple languages at once.</P><PRE><CODE>bdc:WorkCenter a owl:Class ; rdfs:label "WorkCenter"@en , "작업장"@ko ; rdfs:comment "Production work center."@en , "생산 작업장."@ko .</CODE></PRE><P class="">In SPARQL,<SPAN>&nbsp;</SPAN><CODE>FILTER(LANG(?label) = "ko")</CODE><SPAN>&nbsp;</SPAN>selects the desired language. This works because RDF stores<SPAN>&nbsp;</SPAN><CODE>"작업장"@ko</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>"WorkCenter"@en</CODE><SPAN>&nbsp;</SPAN>as<SPAN>&nbsp;</SPAN><STRONG>distinct literals</STRONG><SPAN>&nbsp;</SPAN>from the start. The language tag is part of the value itself, and SPARQL's<SPAN>&nbsp;</SPAN><CODE>LANG()</CODE><SPAN>&nbsp;</SPAN>function reads it. This works by default, with no additional configuration required.</P><H4 id="synonyms-reusing-skosaltlabel" id="toc-hId--780335772">Synonyms: reusing skos:altLabel</H4><P class=""><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>provides a human-readable name, but it does not distinguish between preferred and alternative names. In natural language search or GraphRAG contexts, the same node must be discoverable whether the user searches for<SPAN>&nbsp;</SPAN><EM>"supplier"</EM><SPAN>&nbsp;</SPAN>or<SPAN>&nbsp;</SPAN><EM>"vendor."</EM><SPAN>&nbsp;</SPAN>For this, a way to represent synonyms is required.</P><P class="">This is where we reuse<SPAN>&nbsp;</SPAN><STRONG>SKOS</STRONG><SPAN>&nbsp;</SPAN>(<EM>Simple Knowledge Organization System</EM>, W3C 2009 Recommendation) and its<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE>. SKOS is originally a standard for classification systems and controlled vocabularies. Its central concept,<SPAN>&nbsp;</SPAN><CODE>skos:Concept</CODE>, represents<SPAN>&nbsp;</SPAN><EM>"a unit of thought,"</EM><SPAN>&nbsp;such as&nbsp;</SPAN>a classification value or a category item. This is different from an OWL Class or instance such as<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter</CODE><SPAN>&nbsp;</SPAN>in our KG.</P><P class="">The three SKOS labels:</P><UL class=""><LI><CODE>skos:prefLabel</CODE>: preferred form (one per language is recommended)</LI><LI><CODE>skos:altLabel</CODE>: synonyms, abbreviations, and notation variants (no count limit)</LI><LI><CODE>skos:hiddenLabel</CODE>: labels indexed for search but not displayed (typos, deprecated forms, and so on)</LI></UL><P class="">All three are<SPAN>&nbsp;</SPAN><STRONG>sub-properties</STRONG><SPAN>&nbsp;</SPAN>(<CODE>rdfs:subPropertyOf</CODE>) of<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE>. That is, attaching<SPAN>&nbsp;</SPAN><CODE>prefLabel</CODE><SPAN>&nbsp;</SPAN>already implies<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE>.</P><P class=""><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>does not enforce<SPAN>&nbsp;</SPAN><CODE>rdfs:domain</CODE><SPAN>&nbsp;</SPAN>of<SPAN>&nbsp;</SPAN><CODE>skos:Concept</CODE>. The SKOS specification intentionally omits a domain so that altLabel can be attached to resources of any type.<SPAN>&nbsp;<STRONG>U</STRONG></SPAN><STRONG>sing<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;on a class or a <CODE>skos:Concept</CODE>, is a borrowed use of the property</SPAN></STRONG>.</P><TABLE><TBODY><TR><TD><STRONG>Target</STRONG></TD><TD><STRONG>Meaning</STRONG></TD><TD><STRONG>Example</STRONG></TD><TD><STRONG>skos:altLabel</STRONG></TD></TR><TR><TD><CODE>skos:Concept</CODE></TD><TD>A classification value being referenced</TD><TD>Categories such as<SPAN>&nbsp;</SPAN><EM>"Domestic Supplier,"</EM><SPAN>&nbsp;</SPAN><EM>"Overseas Supplier"</EM></TD><TD><STRONG>Native use</STRONG></TD></TR><TR><TD><CODE>owl:Class</CODE></TD><TD>A set with members</TD><TD><EM>"There is a kind called Supplier"</EM></TD><TD><STRONG>Borrowed</STRONG></TD></TR><TR><TD>Instance</TD><TD>An actual entity</TD><TD>Real Suppliers such as<SPAN>&nbsp;</SPAN><EM>"USSU-VSF01,"</EM><SPAN>&nbsp;</SPAN><EM>"USSU-VSF02"</EM></TD><TD><STRONG>Borrowed</STRONG></TD></TR></TBODY></TABLE><PRE><CODE>bdc:Supplier (Class) ├── USSU-VSF01 → hasRole → c_domestic (Domestic Supplier) └── USSU-VSF02 → hasRole → c_overseas (Overseas Supplier) c_domestic, c_overseas → skos:Concept (classification values)</CODE></PRE><BLOCKQUOTE dir="auto"><P class="">Domestic/overseas classification can also be expressed via a DatatypeProperty.<SPAN>&nbsp;</SPAN><CODE>skos:Concept</CODE><SPAN> is used when the classification value itself must be modeled as an independent concept, for example when it is shared across multiple systems or requires its own label or description</SPAN>.</P></BLOCKQUOTE><P class=""><STRONG>Borrowing on a class (the approach used in this project):</STRONG></P><PRE><CODE>@prefix skos: &lt;http://www.w3.org/2004/02/skos/core#&gt; . bdc:Supplier a owl:Class ; # owl:Class — borrowed use rdfs:label "Supplier"@en , "공급사"@ko ; # preferred name (for display) skos:altLabel "공급업체"@ko , # synonym (for search) "벤더"@ko , "Vendor"@en ; rdfs:comment "A counterparty that supplies materials or services."@en .</CODE></PRE><P class=""><STRONG>The same pattern applies to instances:</STRONG></P><PRE><CODE>ex:successFactors a ex:Product ; # Instance — also borrowed rdfs:label "SAP SuccessFactors"@en ; skos:altLabel "SuccessFactors"@en ; skos:altLabel "SFSF"@en ; skos:altLabel "석세스팩터스"@ko . # Korean notation</CODE></PRE><P class=""><STRONG>In this project, the pattern is applied only at the class level.</STRONG><SPAN>&nbsp;</SPAN>Instance-level synonyms would require leveraging source columns (Virtual Table) or building a separate mapping, which is outside the scope of these examples.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>The line not to cross is this</STRONG>: if synonyms are needed, do not change the resource type to <CODE>skos:Concept</CODE>. That would be a misuse. <CODE>bdc:Supplier</CODE><SPAN>&nbsp;</SPAN>remains an<SPAN>&nbsp;</SPAN><CODE>owl:Class</CODE>,<SPAN>&nbsp;</SPAN><CODE>ex:successFactors</CODE> remains an instance, and only <CODE>skos:altLabel</CODE> is layered on top.</P></BLOCKQUOTE><P class="">To search across the preferred name and synonyms in a single SPARQL query, use property path alternation (<CODE>|</CODE>) :</P><PRE><CODE>SELECT ?class ?label WHERE { ?class a owl:Class . ?class rdfs:label|skos:altLabel ?label . FILTER(CONTAINS(LCASE(?label), LCASE("vendor"))) }</CODE></PRE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":open_book:">📖</span><SPAN>&nbsp;</SPAN><STRONG>Standard basis:</STRONG><SPAN>&nbsp;</SPAN>The SKOS Reference (W3C Recommendation, 2009) does not declare a domain on<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE>, and presents graphs attaching labels to<SPAN>&nbsp;</SPAN><CODE>owl:Class</CODE><SPAN>&nbsp;</SPAN>as official examples. This is not an informal workaround but a legitimate use that the standard explicitly permits.<SPAN>&nbsp;</SPAN><A href="https://www.w3.org/TR/skos-reference/" target="_blank" rel="noopener nofollow noreferrer">https://www.w3.org/TR/skos-reference/</A></P></BLOCKQUOTE><H4 id="can-synonyms-work-without-skos" id="toc-hId--976849277">Can synonyms work without SKOS?</H4><P class="">Yes.<SPAN>&nbsp;</SPAN><STRONG>A synonym is ultimately just a triple, and can be expressed without SKOS.</STRONG><SPAN>&nbsp;</SPAN>The important point is that SKOS does not<SPAN>&nbsp;</SPAN><EM>"enable"</EM><SPAN>&nbsp;</SPAN>synonyms in the first place.</P><P class="">The simplest approach is to define a custom property:</P><PRE><CODE>ex:successFactors a ex:Product ; rdfs:label "SAP SuccessFactors"@en ; ex:alias "SFSF"@en, "SuccessFactors"@en, "석세스팩터스"@ko .</CODE></PRE><P class=""><CODE>ex:alias</CODE><SPAN>&nbsp;</SPAN>is a custom property. Search works exactly the same way (<CODE>?p rdfs:label|ex:alias ?label</CODE>). For GraphRAG-style search alone, this is sufficient.</P><P class="">Alternatively, all variants can be packed into<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE>, and search will still work. The trade-off is that the distinction between preferred and alternative forms disappears, so<SPAN>&nbsp;</SPAN><EM>"which one to show as the primary form in a UI"</EM><SPAN>&nbsp;</SPAN>becomes ambiguous.</P><P class="">So what does SKOS provide<SPAN>&nbsp;</SPAN><EM>in addition</EM>? All three approaches deliver the same<SPAN>&nbsp;</SPAN><EM>"synonym search."</EM><SPAN>&nbsp;</SPAN>The value of SKOS lies elsewhere:</P><P>&nbsp;</P><TABLE><TBODY><TR><TD>&nbsp;</TD><TD><STRONG><CODE>ex:alias</CODE>(custom)&nbsp;&nbsp;</STRONG><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span></TD><TD><STRONG><CODE>rdfs:label</CODE>packed</STRONG><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span></TD><TD><STRONG><CODE>skos:altLabel</CODE></STRONG><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span></TD></TR><TR><TD>Synonym search</TD><TD><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span></TD><TD><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span></TD><TD><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span></TD></TR><TR><TD>Preferred/alternative<SPAN>&nbsp;</SPAN><STRONG>role distinction</STRONG></TD><TD>Build it yourself</TD><TD><span class="lia-unicode-emoji" title=":cross_mark:">❌</span></TD><TD><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span>built in (pref/alt/hidden)</TD></TR><TR><TD><STRONG>Standard</STRONG><SPAN>&nbsp;</SPAN>(recognized by tools and people)</TD><TD><span class="lia-unicode-emoji" title=":cross_mark:">❌</span>custom property</TD><TD>△</TD><TD><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span>W3C</TD></TR><TR><TD>Cross-system<SPAN>&nbsp;</SPAN><STRONG>alignment anchor</STRONG></TD><TD><span class="lia-unicode-emoji" title=":cross_mark:">❌</span></TD><TD><span class="lia-unicode-emoji" title=":cross_mark:">❌</span></TD><TD><span class="lia-unicode-emoji" title=":heavy_large_circle:">⭕</span>(mapping properties)</TD></TR></TBODY></TABLE><P class="">Borrowing<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>is not purely about future-proofing. Its value spans two time horizons:</P><UL class=""><LI><STRONG>Immediate value</STRONG>: as a standard, other tools and people immediately recognize it as<SPAN>&nbsp;</SPAN><EM>"a synonym,"</EM><SPAN>&nbsp;</SPAN>and the role distinction (pref/alt/hidden) is predefined, so there is no need to reinvent the wheel.</LI><LI><STRONG>Future value</STRONG>: when extending into cross-system mapping (<CODE>skos:exactMatch</CODE><SPAN>&nbsp;</SPAN>and similar), the same vocabulary family is reused.</LI></UL><H4 id="when-to-use-what" id="toc-hId--1173362782">When to use what</H4><TABLE><TBODY><TR><TD><STRONG>What you need </STRONG></TD><TD><STRONG>Approach</STRONG></TD></TR><TR><TD>One preferred name</TD><TD><CODE>rdfs:label</CODE></TD></TR><TR><TD>Synonyms, abbreviations</TD><TD><CODE>skos:altLabel</CODE></TD></TR><TR><TD>Multilingual preferred name</TD><TD><CODE>rdfs:label "이름"@ko , "Name"@en</CODE></TD></TR><TR><TD>Multilingual synonyms</TD><TD><CODE>skos:altLabel "벤더"@ko , "Vendor"@en</CODE></TD></TR><TR><TD>Classification or code system</TD><TD><CODE>skos:Concept</CODE><SPAN>&nbsp;</SPAN>+<SPAN>&nbsp;</SPAN><CODE>skos:prefLabel</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN><EM>(out of scope for this example; extend when needed)</EM></TD></TR></TBODY></TABLE><P class="">Typical situations where synonyms must be attached:</P><UL class=""><LI><STRONG>GraphRAG or natural language search</STRONG>: the same node must be discoverable regardless of how the user phrases the query.</LI><LI><STRONG>Multilingual environments</STRONG>: a single concept must be searchable in both English and Korean.</LI><LI><STRONG>System integration</STRONG>: terms from other systems must be marked as representing the same concept.</LI></UL><HR /><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":open_book:">📖</span><SPAN>&nbsp;</SPAN><STRONG>Reference:</STRONG><SPAN>&nbsp;</SPAN>The concepts covered in this section (Competency Questions, Class vs Instance, Domain &amp; Range) are based on<SPAN>&nbsp;</SPAN><STRONG>Ontology Development 101</STRONG><SPAN>&nbsp;</SPAN>(Noy &amp; McGuinness, Stanford 2001), the standard introduction to ontology design.<SPAN>&nbsp;</SPAN><A href="https://protege.stanford.edu/publications/ontology_development/ontology101.pdf" target="_blank" rel="noopener nofollow noreferrer">https://protege.stanford.edu/publications/ontology_development/ontology101.pdf</A></P></BLOCKQUOTE><HR /><H1 id="2-the-seven-steps-of-ontology-101-the-process-we-followed" id="toc-hId--489667266">2. The Seven Steps of Ontology 101: The Process We Followed</H1><P class="">The concepts covered in Section 1 (Class/Instance, Domain &amp; Range, Competency Questions) all originate from Stanford's<SPAN>&nbsp;</SPAN><EM>Ontology Development 101</EM>. This guide does not just enumerate concepts. It proposes a<SPAN>&nbsp;</SPAN><STRONG>seven-step process for building an ontology</STRONG>.</P><P class="">We followed this seven-step process directly when designing the BDC Data Products-based KG. Before going into the practical work in the following sections, here is the high-level mapping.</P><H3 id="the-seven-step-process" id="toc-hId--1272986785">The seven-step process</H3><TABLE><TBODY><TR><TD><STRONG>Step</STRONG></TD><TD><STRONG>Original step </STRONG></TD><TD><STRONG>Meaning</STRONG></TD><TD><STRONG><EM>Mapping in our project</EM></STRONG></TD></TR><TR><TD>1</TD><TD>Determine the domain and scope</TD><TD>Determine domain and scope (with<SPAN>&nbsp;</SPAN><STRONG>Competency Questions</STRONG>)</TD><TD><EM>"How far does the impact spread when Z_ASM3 stops?"</EM><SPAN>&nbsp;</SPAN>(the four CQs from Section 1)</TD></TR><TR><TD>2</TD><TD>Consider reusing existing ontologies</TD><TD>Evaluate reuse of existing ontologies</TD><TD><STRONG>One Domain Model</STRONG>: standard vocabulary (<CODE>WorkCenter</CODE>,<SPAN>&nbsp;</SPAN><CODE>Supplier</CODE>,<SPAN>&nbsp;</SPAN><CODE>Product</CODE>) used as-is</TD></TR><TR><TD>3</TD><TD>Enumerate important terms</TD><TD>Enumerate important terms</TD><TD>Exploring entities and column names from BDC Virtual Tables (Part 3 Section 1)</TD></TR><TR><TD>4</TD><TD>Define classes and class hierarchy</TD><TD>Define classes and the class hierarchy</TD><TD><CODE>MasterData</CODE>/<CODE>StructuralNode</CODE><SPAN>&nbsp;</SPAN>upper layer + six classes (including the Product partition in Part 2-2)</TD></TR><TR><TD>5</TD><TD>Define properties of classes</TD><TD>Define the properties of classes</TD><TD><CODE>producesProduct</CODE>,<SPAN>&nbsp;</SPAN><CODE>locatedAt</CODE>,<SPAN>&nbsp;</SPAN><CODE>hasComponent</CODE>, and other ObjectProperties</TD></TR><TR><TD>6</TD><TD>Define facets of properties</TD><TD>Define property facets (cardinality, value type)</TD><TD>Domain/Range constraints,<SPAN>&nbsp;</SPAN><CODE>owl:disjointWith</CODE><SPAN>&nbsp;</SPAN>(Part 2-2)</TD></TR><TR><TD>7</TD><TD>Create instances</TD><TD>Create instances</TD><TD>Part 3 Section 5 (Relationship Views → TTL triples)</TD></TR></TBODY></TABLE><H3 id="if-you-remember-one-thing-iteration-is-the-essence" id="toc-hId--1469500290">If you remember one thing: iteration is the essence</H3><P class="">The point that Ontology 101 emphasizes most is that<SPAN>&nbsp;</SPAN><STRONG>"it is never completed in one pass."</STRONG></P><BLOCKQUOTE dir="auto"><P class=""><EM>"Designing an ontology is necessarily an iterative process."</EM></P></BLOCKQUOTE><P class="">Attempting to model an entire domain in one pass is rarely realistic:</P><UL class=""><LI><STRONG>Scope is hard to bound</STRONG>: the boundary of<SPAN>&nbsp;</SPAN><EM>"how far to model"</EM><SPAN>&nbsp;</SPAN>becomes unclear.</LI><LI><STRONG>Completion is hard to verify</STRONG>: without a defined test scope, the quality benchmark is unclear.</LI><LI><STRONG>Project milestones are hard to set</STRONG>: schedules and deliverables blur, making productization difficult.</LI></UL><P class="">After one full pass through the seven steps, you have the minimum KG required to answer the first CQ. When a new CQ arrives, you return to Step 1 and run the cycle again.</P><P class="">To borrow the book's phrasing:</P><BLOCKQUOTE dir="auto"><P class=""><EM>"Develop the ontology incrementally, one competency question at a time."</EM></P></BLOCKQUOTE><P class=""><STRONG>Class Partition (Part 2-2), the Relationship SQL View (Part 3 Section 3), and the decision to leave Customer in SQL (Part 2-2)</STRONG><SPAN>&nbsp;</SPAN>were not produced in the first iteration. They were solutions to problems discovered while validating the CQs.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":open_book:">📖</span><SPAN>&nbsp;</SPAN><STRONG>Source:</STRONG><SPAN>&nbsp;</SPAN>Noy, N. F., &amp; McGuinness, D. L. (2001).<SPAN>&nbsp;</SPAN><EM>Ontology Development 101: A Guide to Creating Your First Ontology</EM>. Stanford KSL Technical Report.<SPAN>&nbsp;</SPAN><A href="https://protege.stanford.edu/publications/ontology_development/ontology101.pdf" target="_blank" rel="noopener nofollow noreferrer">https://protege.stanford.edu/publications/ontology_development/ontology101.pdf</A></P></BLOCKQUOTE><HR /><H1 id="3-from-rdb-to-rdf-what-the-w3c-standards-define" id="toc-hId--1079207781">3. From RDB to RDF: What the W3C Standards Define</H1><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":information:">ℹ️</span><SPAN>&nbsp;</SPAN><STRONG>Where this section sits:</STRONG><SPAN>&nbsp;</SPAN>If Section 1 and Section 2 dealt with the ontology (schema) itself, this section covers its counterpart, the<SPAN>&nbsp;</SPAN><STRONG>instance (data) translation standard</STRONG>. The two together make a KG. What Section 2 calls Step 7 (Create Instances) in the Ontology 101 process is exactly what the W3C standardized here.</P></BLOCKQUOTE><P class="">So far we have covered<SPAN>&nbsp;</SPAN><EM>"what an ontology is"</EM><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><EM>"how to design one."</EM><SPAN>&nbsp;</SPAN>When data already lives in a relational database, as with SAP BDC, the question of how to translate it into RDF triples was standardized by the W3C in 2012.</P><P class="">The automatic translation by the standard alone is not sufficient. SAP data often involves business relationships that span multiple tables, that exist without FK declarations, or that require a single table to be split into multiple classes. A more flexible approach is needed. For this reason, the W3C standardized both automatic translation (Direct Mapping) and custom mapping (R2RML) together.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":triangular_ruler:">📐</span><SPAN>&nbsp;</SPAN><STRONG>R2RML produces instance data, not ontology</STRONG></P><P class="">R2RML is frequently misunderstood at first encounter.<SPAN>&nbsp;</SPAN><EM>"Does this produce an ontology, or instance data?"</EM><SPAN>&nbsp;</SPAN>The answer is<SPAN>&nbsp;</SPAN><STRONG>instance data</STRONG>. R2RML is a file that contains only<SPAN>&nbsp;</SPAN><STRONG>transformation rules</STRONG><SPAN>&nbsp;</SPAN>such as<SPAN>&nbsp;</SPAN><EM>"map this row of this table as an instance of which class, and map this column as which property."</EM><SPAN>&nbsp;</SPAN>The class definitions themselves (<CODE>bdc:Supplier is a class with the following properties</CODE>) live in a separate ontology file, and R2RML simply<SPAN>&nbsp;</SPAN><EM>references</EM><SPAN>&nbsp;</SPAN>them.</P><P class="">Three files form a useful mental model:</P><TABLE><TBODY><TR><TD><STRONG>File</STRONG></TD><TD><STRONG>What it contains</STRONG></TD><TD><STRONG>Example</STRONG></TD></TR><TR><TD><STRONG>Ontology file</STRONG></TD><TD>Class and property definitions (the blueprint)</TD><TD><CODE>bdc:Supplier a owl:Class</CODE></TD></TR><TR><TD><STRONG>R2RML mapping file</STRONG></TD><TD>The rule<SPAN>&nbsp;</SPAN><EM>"take a row from this table (or the result of this SQL query) and convert it to instance data according to that blueprint"</EM></TD><TD><CODE>Row of SUPPLIER_MASTER = instance of bdc:Supplier</CODE></TD></TR><TR><TD><STRONG>Instance data</STRONG></TD><TD>The actual data triples</TD><TD><CODE>bdc:Supplier_S001 a bdc:Supplier ; bdc:name "ACME"</CODE></TD></TR></TBODY></TABLE></BLOCKQUOTE><P class="">There are two standards:</P><P>&nbsp;</P><TABLE><TBODY><TR><TD><STRONG>&nbsp;</STRONG></TD><TD><STRONG>Direct Mapping </STRONG></TD><TD><STRONG>R2RML</STRONG></TD></TR><TR><TD><STRONG>Translation method</STRONG></TD><TD>Automatic, based on schema</TD><TD>Mapping file authored manually</TD></TR><TR><TD><STRONG>Class / relationship / attribute names</STRONG></TD><TD>Table and column names used as-is (<CODE>Person_table</CODE>,<SPAN>&nbsp;</SPAN><CODE>Person_table#name</CODE>)</TD><TD>Free to assign business vocabulary (<CODE>bdc:Supplier</CODE>,<SPAN>&nbsp;</SPAN><CODE>bdc:hasComponent</CODE>,<SPAN>&nbsp;</SPAN><CODE>bdc:supplierName</CODE>)</TD></TR><TR><TD><STRONG>When to use</STRONG></TD><TD>When every table maps cleanly to a single class and FKs accurately express all business relationships</TD><TD>When one table needs to be split into multiple classes (for example, splitting<SPAN>&nbsp;</SPAN><CODE>MaterialType</CODE><SPAN>&nbsp;</SPAN>into FinishedGood/SemiFinished/RawMaterial), when FKs are missing or relationships span multi-hop chains, or when ABAP abbreviations (<CODE>EKKO</CODE>,<SPAN>&nbsp;</SPAN><CODE>KNA1</CODE>) need to be replaced with business vocabulary (<CODE>PurchaseOrder</CODE>,<SPAN>&nbsp;</SPAN><CODE>Customer</CODE>)</TD></TR></TBODY></TABLE><P class=""><STRONG>Direct Mapping provides the baseline automatic mapping from relational data to RDF, while R2RML provides a customizable mapping language for cases where the default mapping is not sufficient.</STRONG><SPAN>&nbsp;</SPAN>Below, each of the four mapping types is presented side by side as<SPAN>&nbsp;</SPAN><EM>"what Direct Mapping does automatically"</EM><SPAN>&nbsp;</SPAN>versus<SPAN>&nbsp;</SPAN><EM>"what R2RML / our project actually did."</EM></P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":memo:">📝</span><SPAN>&nbsp;</SPAN><STRONG>Note:</STRONG><SPAN>&nbsp;</SPAN>Our project did not use an R2RML engine directly. We generated triples using Python scripts, but the principles are identical to those defined by R2RML.</P></BLOCKQUOTE><H3 id="a-preview-of-r2rml-standard-vocabulary" id="toc-hId--1862527300">A preview of R2RML standard vocabulary</H3><P class="">Sections 3.1–3.4 below cover the four mapping types. Before entering the detail, here is a preview of how our SQL View and Python approach maps to R2RML standard vocabulary, limited to the three core constructs. As you read Sections 3.1–3.4, recognizing<SPAN>&nbsp;</SPAN><EM>"this corresponds to that R2RML construct"</EM><SPAN>&nbsp;</SPAN>is sufficient.</P><TABLE><TBODY><TR><TD><STRONG>R2RML construct</STRONG></TD><TD><STRONG>What the standard defines</STRONG></TD><TD><STRONG>Step in Section 3</STRONG></TD></TR><TR><TD><CODE>rr:tableName</CODE></TD><TD>Designates a single table as the mapping input source</TD><TD>Section 3.1: single table → single class</TD></TR><TR><TD><CODE>rr:sqlQuery</CODE></TD><TD>Designates an arbitrary SQL SELECT result as the mapping input source</TD><TD>Section 3.1: multiple tables JOINed → single class (for example, SUPPLIER_MASTER + DETAIL →<SPAN>&nbsp;</SPAN><CODE>bdc:Supplier</CODE>) / Section 3.4: multi-hop relationship View</TD></TR><TR><TD><CODE>rr:joinCondition</CODE></TD><TD>Joins two tables inside a RefObjectMap to produce relationship triples</TD><TD>Section 3.4: FK join is precomputed inside the View; two IRIs are exposed together in a single row</TD></TR></TBODY></TABLE><P class="">This table pairs only<SPAN>&nbsp;</SPAN><EM>"what (Instance Data)"</EM><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><EM>"how connected (Relationship)."</EM><SPAN>&nbsp;</SPAN>The IRI construction in Section 3.2 Instance Data Identification (<CODE>rr:subjectMap</CODE>) and the column → attribute mapping in Section 3.3 Data Attribute (<CODE>rr:predicateObjectMap</CODE><SPAN>&nbsp;</SPAN>of literal form) are straightforward and not given separate rows.</P><P class="">Inverting this view yields the picture of<SPAN>&nbsp;</SPAN><STRONG>using the DB itself as the mapping engine</STRONG>, without a separate R2RML processing engine. When a View exposes every necessary column in a single row (the subject's ID and the relationship target's ID), the JOIN that an R2RML engine would perform inside its mapping is already done by the database. The output (relationship triples) is the same.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":pushpin:">📌</span><SPAN>&nbsp;</SPAN><STRONG>The<SPAN>&nbsp;</SPAN><EM>Relationship SQL View</EM><SPAN>&nbsp;</SPAN>pattern in Part 3 is precisely this picture.</STRONG><SPAN>&nbsp;</SPAN>The HANA SQL Views (<CODE>KGR_*</CODE>) precompute multi-hop JOINs, and the TTL generation script reads those Views to write relationship triples directly. The transformation semantics defined by the standard are followed exactly, without a separate mapping engine.</P></BLOCKQUOTE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":information:">ℹ️</span><SPAN>&nbsp;</SPAN><STRONG>Note: another mode.</STRONG><SPAN>&nbsp;</SPAN>The discussion above assumes the<SPAN>&nbsp;</SPAN><EM>"materialize relationship triples ahead of time and write them into the graph"</EM><SPAN>&nbsp;</SPAN>approach. The standard defines another mode as well. Instead of materializing triples in advance, joinConditions are translated into SQL JOINs on the fly each time a SPARQL query arrives. This series follows the materialized approach.</P></BLOCKQUOTE><HR /><H2 id="31-class-definition" id="toc-hId--1597454107">3.1 Class definition</H2><H3 id="w3c-direct-mapping" id="toc-hId--2087370619">W3C Direct Mapping</H3><P class="">One table becomes one class. Each row is automatically translated into an instance triple of that class.</P><PRE><CODE>&lt;People/ID=7&gt; rdf:type &lt;People&gt; .</CODE></PRE><P class="">However,<SPAN>&nbsp;</SPAN><STRONG>the class definition itself is not generated.</STRONG><SPAN>&nbsp;</SPAN>Direct Mapping outputs only instance triples. The declaration that<SPAN>&nbsp;</SPAN><CODE>&lt;People&gt;</CODE><SPAN>&nbsp;</SPAN>is a class must be written separately.</P><PRE><CODE>&lt;People&gt; rdf:type rdfs:Class ; rdfs:label "Person" ; rdfs:comment "Business entity representing a person" .</CODE></PRE><H3 id="r2rml--our-project" id="toc-hId-2011083172">R2RML / our project</H3><P class="">Direct Mapping's<SPAN>&nbsp;</SPAN><EM>"1 Table = 1 Class"</EM><SPAN>&nbsp;</SPAN>fits simple relational databases well, but real SAP data is different.<SPAN>&nbsp;</SPAN><STRONG>Class and Table are not a 1:1 correspondence.</STRONG></P><P class=""><STRONG>1 Table → N Classes (partition):</STRONG><SPAN>&nbsp;</SPAN>when a single table contains multiple kinds of entities. The BDC Product table is split into three classes based on the<SPAN>&nbsp;</SPAN><CODE>MaterialType</CODE><SPAN>&nbsp;</SPAN>column.</P><PRE><CODE>BDC_DP.PRODUCT (single table) MaterialType = FERT → bdc:FinishedGood MaterialType = HALB → bdc:SemiFinished MaterialType = ROH → bdc:RawMaterial</CODE></PRE><P class=""><STRONG>N Tables → 1 Class (consolidation):</STRONG><SPAN>&nbsp;</SPAN>when multiple tables form a single business concept.<SPAN>&nbsp;</SPAN><CODE>WorkCenter</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>WorkCenterCapacity</CODE><SPAN>&nbsp;</SPAN>are separate tables, but in the KG they are represented by a single class<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter</CODE>.</P><P class=""><EM>"Which data belongs to which class"</EM><SPAN>&nbsp;</SPAN>is determined by Competency Questions and domain decisions, not by table schema. Class Partition is covered in detail in Part 2-2.</P><P class="">In our project,<SPAN>&nbsp;</SPAN><CODE>bdc_dp_ontology.ttl</CODE><SPAN>&nbsp;</SPAN>is the file that holds the class definitions.</P><HR /><H2 id="32-instance-data-identification" id="toc-hId-2107972674">3.2 Instance data identification</H2><H3 id="w3c-direct-mapping-1" id="toc-hId-1618056162">W3C Direct Mapping</H3><P class="">Each row becomes an<SPAN>&nbsp;</SPAN><STRONG>instance (node)</STRONG><SPAN>&nbsp;</SPAN>in the KG. It is identified by an IRI that combines the table name and the primary key value.</P><PRE><CODE>&lt;People/ID=7&gt; # single PK &lt;Order/customer=7;item=42&gt; # composite PK separated by ;</CODE></PRE><H3 id="r2rml--our-project-1" id="toc-hId-1421542657">R2RML / our project</H3><P class="">The IRI format itself is free. The standard simply provides a guide for<SPAN>&nbsp;</SPAN><EM>"how names are constructed during automatic translation."</EM><SPAN>&nbsp;</SPAN>We define readable forms such as<SPAN>&nbsp;</SPAN><CODE>bdc:Person_7</CODE><SPAN>&nbsp;</SPAN>directly.</P><P class="">In our project, each row of a BDC Virtual Table generates an instance IRI in this pattern.</P><PRE><CODE>bdc:Person_7 a bdc:Person ; rdfs:label "Bob" .</CODE></PRE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span><SPAN>&nbsp;</SPAN><STRONG>Always attach<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>to instances.</STRONG><SPAN>&nbsp;</SPAN>IRIs are identifiers, not human-readable names. SPARQL results returning only IRIs are difficult for an LLM to interpret.</P></BLOCKQUOTE><HR /><H2 id="33-data-attribute-representation" id="toc-hId-1518432159">3.3 Data attribute representation</H2><H3 id="w3c-direct-mapping-2" id="toc-hId-1028515647">W3C Direct Mapping</H3><P class="">Each column of a table is translated into a literal property. The predicate name is automatically generated in the<SPAN>&nbsp;</SPAN><CODE>tableName#columnName</CODE><SPAN>&nbsp;</SPAN>form. For example, the<SPAN>&nbsp;</SPAN><CODE>fname</CODE><SPAN>&nbsp;</SPAN>column of the People table becomes<SPAN>&nbsp;</SPAN><CODE>&lt;People#fname&gt;</CODE>.</P><PRE><CODE>&lt;People/ID=7&gt; &lt;People#fname&gt; "Bob" . &lt;People/ID=7&gt; &lt;People#ID&gt; 7 .</CODE></PRE><P class="">NULL values do not produce triples.</P><H3 id="r2rml--our-project-2" id="toc-hId-832002142">R2RML / our project</H3><P class="">Predicate names can be replaced with business vocabulary.<SPAN>&nbsp;</SPAN><CODE>fname</CODE><SPAN>&nbsp;</SPAN>is the column name, but the business term<SPAN>&nbsp;</SPAN><CODE>name</CODE><SPAN>&nbsp;</SPAN>is more appropriate, so we define<SPAN>&nbsp;</SPAN><CODE>bdc:name</CODE><SPAN>&nbsp;</SPAN>directly. Property definitions (domain, range, label) must be written separately in the ontology TTL.</P><PRE><CODE># Defined in ontology TTL bdc:name rdf:type owl:DatatypeProperty ; rdfs:domain bdc:Person ; rdfs:range xsd:string ; rdfs:label "Name" . # Used in instance TTL bdc:Person_7 bdc:name "Bob" .</CODE></PRE><HR /><H2 id="34-relationship-representation" id="toc-hId-928891644">3.4 Relationship representation</H2><H3 id="w3c-direct-mapping-3" id="toc-hId-607158823">W3C Direct Mapping</H3><P class="">A single FK column produces two triples.</P><PRE><CODE>&lt;People/ID=7&gt; &lt;People#addressID&gt; 18 . # the value itself (literal triple) &lt;People/ID=7&gt; &lt;People#ref-addressID&gt; &lt;Addresses/ID=18&gt; . # the reference (reference triple)</CODE></PRE><UL class=""><LI><STRONG>literal triple</STRONG>: preserves the original FK value.</LI><LI><STRONG>reference triple</STRONG>: a graph edge pointing to another instance.</LI></UL><P class="">However, the predicate name is fixed in the<SPAN>&nbsp;</SPAN><CODE>#ref-colName</CODE><SPAN>&nbsp;</SPAN>form and cannot carry business meaning. Furthermore,<SPAN>&nbsp;</SPAN><STRONG>Direct Mapping creates a reference triple only when an FK is declared.</STRONG></P><H3 id="r2rml--our-project-3" id="toc-hId-410645318">R2RML / our project</H3><P class="">R2RML does not require FK declarations.<SPAN>&nbsp;</SPAN><STRONG>As long as the values exist</STRONG>, a relationship can be created.</P><BLOCKQUOTE dir="auto"><P class="">If FKs cover 100% of relationships, Direct Mapping alone may be sufficient. In real SAP data, two reasons make this rarely the case.</P><P class="">First, once classes are split, the relationships between those classes must be defined explicitly. When<SPAN>&nbsp;</SPAN><CODE>MaterialType</CODE><SPAN>&nbsp;</SPAN>partitions Products into<SPAN>&nbsp;</SPAN><CODE>FinishedGood</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>SemiFinished</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>RawMaterial</CODE>, the relationships among them (such as<SPAN>&nbsp;</SPAN><CODE>hasComponent</CODE>) are determined by domain analysis, not by FKs.</P><P class="">Second, when a business relationship spans multiple tables (multi-hop), a single FK cannot connect them directly. This is why we precompute relationships with SQL Views to flatten them into 1-hop edges.</P><P class="">Ultimately, relationships are determined by<SPAN>&nbsp;</SPAN><STRONG>people who know the domain</STRONG>, not by FKs.</P></BLOCKQUOTE><P class="">There are three approaches to creating relationships:</P><P class=""><STRONG>Approach A: Assemble the IRI directly from a value</STRONG><SPAN>&nbsp;</SPAN>(simple case)</P><P class="">Take the value of the FK column and construct an IRI directly. There is no need to look at the neighboring table. If<SPAN>&nbsp;</SPAN><CODE>addressID</CODE><SPAN>&nbsp;</SPAN>is<SPAN>&nbsp;</SPAN><CODE>18</CODE>, construct the IRI<SPAN>&nbsp;</SPAN><CODE>bdc:Address_18</CODE><SPAN>&nbsp;</SPAN>and create the relationship triple.</P><PRE><CODE>bdc:Person_7 bdc:hasAddress bdc:Address_18 .</CODE></PRE><P class=""><STRONG>Approach B: Precompute the relationship with a SQL View</STRONG><SPAN>&nbsp;</SPAN>(complex case)</P><P class="">For multi-hop or complex JOIN relationships, we precompute with a SQL View. The View returns<SPAN>&nbsp;</SPAN><CODE>(subj_id, obj_id)</CODE><SPAN>&nbsp;</SPAN>pairs, and Python reads them to build triples.</P><PRE><CODE>-- KGR_WC_PRODUCES_PRODUCT View result (example) subj_id | obj_id -------------|---------- Z_ASM3 | MZ-FG-C900 Z_ASM3 | MZ-FG-C950</CODE></PRE><P class="">→<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter_Z_ASM3 bdc:producesProduct bdc:FinishedGood_MZ-FG-C900</CODE></P><P class="">Our project's<SPAN>&nbsp;</SPAN><CODE>KGR_WC_LOCATEDAT_PLANT</CODE>,<SPAN>&nbsp;</SPAN><CODE>KGR_WC_PRODUCES_PRODUCT</CODE>, and others use this approach.</P><P class=""><STRONG>Approach C: Connect by shared column, even without FK</STRONG><SPAN>&nbsp;</SPAN>(simple case)</P><P class="">Even without an FK constraint, if both tables share a column with the same meaning, Approach A applies directly. For example, the<SPAN>&nbsp;</SPAN><CODE>People</CODE><SPAN>&nbsp;</SPAN>table and the<SPAN>&nbsp;</SPAN><CODE>Addresses</CODE><SPAN>&nbsp;</SPAN>table are linked by<SPAN>&nbsp;</SPAN><CODE>addressID</CODE>, but no FK is declared:</P><PRE><CODE>People table Addresses table ID=7, addressID=18 → ID=18, city=Seoul ↓ no FK declared Use addressID value 18 to build the IRI → bdc:Address_18</CODE></PRE><P class="">→<SPAN>&nbsp;</SPAN><CODE>bdc:Person_7 bdc:hasAddress bdc:Address_18</CODE></P><P class="">The relationship works as long as the value exists. Our SQL Views also operate independently of whether FKs are declared.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>In practice, all three are used.</STRONG><SPAN>&nbsp;</SPAN>Simple FK relationships use Approach A, connections without FKs use Approach C, and complex multi-hop relationships use Approach B. The right choice depends on the situation, and all three approaches appeared in our project.</P></BLOCKQUOTE><P>&nbsp;</P><TABLE><TBODY><TR><TD>&nbsp;</TD><TD><STRONG>Direct Mapping</STRONG></TD><TD><STRONG>R2RML / our project</STRONG></TD></TR><TR><TD>FK declaration required</TD><TD>Yes</TD><TD>No (values are enough)</TD></TR><TR><TD>Relationship name</TD><TD>Fixed as<SPAN>&nbsp;</SPAN><CODE>#ref-colName</CODE></TD><TD>Free (<CODE>bdc:hasAddress</CODE>, business vocabulary)</TD></TR><TR><TD>Defined by</TD><TD>Machine (W3C spec)</TD><TD>Human (domain expert)</TD></TR><TR><TD>Multi-hop</TD><TD>Not supported (only direct FK)</TD><TD>Supported via precomputed SQL Views, flattened to 1-hop</TD></TR></TBODY></TABLE><HR /><H3 id="combined-example-how-two-tables-become-a-graph" id="toc-hId-214131813">Combined example: how two tables become a graph</H3><P class=""><STRONG>Original SQL tables</STRONG></P><PRE><CODE>-- People table ID | fname | addressID ----+-------+---------- 7 | Bob | 18 -- Addresses table ID | city ----+------ 18 | Seoul</CODE></PRE><P class=""><STRONG>Step 1: Direct Mapping automatic result</STRONG></P><PRE><CODE>&lt;People/ID=7&gt; rdf:type &lt;People&gt; . &lt;Addresses/ID=18&gt; rdf:type &lt;Addresses&gt; . &lt;People/ID=7&gt; &lt;People#fname&gt; "Bob" . &lt;People/ID=7&gt; &lt;People#addressID&gt; 18 . &lt;Addresses/ID=18&gt; &lt;Addresses#city&gt; "Seoul" . &lt;People/ID=7&gt; &lt;People#ref-addressID&gt; &lt;Addresses/ID=18&gt; .</CODE></PRE><P class="">The data has been translated, but the predicates carry technical names (<CODE>&lt;People#fname&gt;</CODE>,<SPAN>&nbsp;</SPAN><CODE>&lt;People#ref-addressID&gt;</CODE>).</P><P class=""><STRONG>Step 2: Layer an ontology to add business meaning</STRONG></P><PRE><CODE># ontology TTL bdc:Person rdf:type owl:Class ; rdfs:label "Person" . bdc:Address rdf:type owl:Class ; rdfs:label "Address" . bdc:hasName rdf:type owl:DatatypeProperty ; rdfs:domain bdc:Person ; rdfs:range xsd:string . bdc:hasAddress rdf:type owl:ObjectProperty ; rdfs:domain bdc:Person ; rdfs:range bdc:Address . bdc:inCity rdf:type owl:DatatypeProperty ; rdfs:domain bdc:Address ; rdfs:range xsd:string . # instance TTL (with business vocabulary) &lt;People/ID=7&gt; rdf:type bdc:Person ; rdfs:label "Bob" ; bdc:hasName "Bob" ; bdc:hasAddress &lt;Addresses/ID=18&gt; . &lt;Addresses/ID=18&gt; rdf:type bdc:Address ; rdfs:label "Seoul" ; bdc:inCity "Seoul" .</CODE></PRE><P class=""><STRONG>The standard's automatic mapping is the starting point; layering ontology definitions to add business meaning is the destination.</STRONG></P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":open_book:">📖</span><SPAN>&nbsp;</SPAN><STRONG>Sources:</STRONG></P><UL class=""><LI>W3C,<SPAN>&nbsp;</SPAN><EM>A Direct Mapping of Relational Data to RDF</EM><SPAN>&nbsp;</SPAN>(Recommendation, 2012-09-27).<SPAN>&nbsp;</SPAN><A href="https://www.w3.org/TR/rdb-direct-mapping/" target="_blank" rel="noopener nofollow noreferrer">https://www.w3.org/TR/rdb-direct-mapping/</A></LI><LI>W3C,<SPAN>&nbsp;</SPAN><EM>R2RML: RDB to RDF Mapping Language</EM><SPAN>&nbsp;</SPAN>(Recommendation, 2012-09-27).<SPAN>&nbsp;</SPAN><A href="https://www.w3.org/TR/r2rml/" target="_blank" rel="noopener nofollow noreferrer">https://www.w3.org/TR/r2rml/</A></LI></UL></BLOCKQUOTE><HR /><H1 id="4-three-domain-decisions-we-made-on-top-of-the-standards" id="toc-hId-604424322">4. Three Domain Decisions We Made on Top of the Standards</H1><P class="">The mappings defined by the W3C standards work well for simple relational databases. For data with<SPAN>&nbsp;</SPAN><STRONG>a rich business domain inside</STRONG>, such as SAP BDC Data Products, a few domain decisions on top of the standard mappings are required to make the KG genuinely useful. We encountered three such points in our project.</P><H3 id="decision-1-class-partition-one-table-multiple-classes" id="toc-hId--178895197">Decision 1. Class Partition: one table, multiple classes</H3><P class="">The default behavior of Direct Mapping is to map one table to one class. This is sufficient for simple RDBs, but in SAP's Product table, finished goods, semi-finished, and raw materials are mixed in the same table and distinguished only by the<SPAN>&nbsp;</SPAN><CODE>MaterialType</CODE><SPAN>&nbsp;</SPAN>column. Mapped as-is, all three would become a single<SPAN>&nbsp;</SPAN><CODE>bdc:Product</CODE>.</P><P class="">→<SPAN>&nbsp;</SPAN><STRONG>Decision: assign different classes per row based on a discriminator column.</STRONG><SPAN>&nbsp;</SPAN>Covered in detail in Part 2-2.</P><H3 id="decision-2-multi-hop-relationships-business-relationships-beyond-direct-fks" id="toc-hId--375408702">Decision 2. Multi-hop relationships: business relationships beyond direct FKs</H3><P class="">Direct Mapping creates reference triples<SPAN>&nbsp;</SPAN><STRONG>only from declared FKs</STRONG>. This has two implications.</P><P class=""><STRONG>(a)</STRONG><SPAN>&nbsp;</SPAN>If an FK constraint is missing from the schema, no relationship triple is produced. The column value remains as a literal, and the two entities are not connected.</P><P class=""><STRONG>(b)</STRONG><SPAN>&nbsp;</SPAN>A relationship requires a direct FK. SAP S/4HANA business relationships are often expressed through<SPAN>&nbsp;</SPAN><STRONG>several intermediate tables</STRONG>. The relationship between WorkCenter and finished goods, for example, spans three routing tables, forming a 4-hop chain. This can be expressed with R2RML join conditions, but we chose a simpler approach.</P><P class="">→<SPAN>&nbsp;</SPAN><STRONG>Decision: precompute multi-hop chains with SQL Views and flatten them into 1-hop relationships.</STRONG><SPAN>&nbsp;</SPAN>Covered in detail in Part 3 Section 3.</P><H3 id="decision-3-junction-tables-keep-as-node-or-flatten" id="toc-hId--571922207">Decision 3. Junction tables: keep as node, or flatten</H3><P class="">N:N relationships in SQL are typically expressed via<SPAN>&nbsp;</SPAN><STRONG>junction tables</STRONG><SPAN>&nbsp;</SPAN>(intermediate tables linking two tables). In our demo,<SPAN>&nbsp;</SPAN><CODE>PurchasingSourceList</CODE><SPAN>&nbsp;</SPAN>(PSL) is such a junction: a list of suppliers approved to provide each product.</P><P class="">When moving such a junction into a KG, there are two options.</P><P class=""><STRONG>Option 1. Keep as a node (Direct Mapping's default behavior):</STRONG><SPAN>&nbsp;</SPAN>PSL becomes a separate node in the graph, with relationships connecting it to both sides.</P><P class=""><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="WorkCenter Product Supply-2026-07-04-065458.svg" style="width: 400px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429108iB1A22A07B45D2181/image-size/medium?v=v2&amp;px=400" role="button" title="WorkCenter Product Supply-2026-07-04-065458.svg" alt="WorkCenter Product Supply-2026-07-04-065458.svg" /></span></P><P class="">→ The PSL node serves as a<SPAN>&nbsp;</SPAN><STRONG>bridge between Supplier and Product</STRONG>. PSL can carry additional attributes such as plant, validity period, and priority.</P><P class=""><STRONG>Option 2. Flatten (direct relationship):</STRONG><SPAN>&nbsp;</SPAN>drop the PSL node and connect Supplier directly to Product.</P><P class=""><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="WorkCenter Product Supply-2026-07-04-065414.svg" style="width: 400px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429107iF29728A824E8B57E/image-size/medium?v=v2&amp;px=400" role="button" title="WorkCenter Product Supply-2026-07-04-065414.svg" alt="WorkCenter Product Supply-2026-07-04-065414.svg" /></span></P><P class="">→ Cleaner, but the additional information PSL carries disappears. For example, if the same supplier-product pair has a different priority per plant, or if expired registrations must be distinguished from current ones, that information is lost. The fact that<SPAN>&nbsp;</SPAN><EM>"Supplier A supplies Product X"</EM><SPAN>&nbsp;</SPAN>remains, but<SPAN>&nbsp;</SPAN><EM>"when, at which plant, with what priority"</EM><SPAN>&nbsp;</SPAN>can no longer be answered by the KG.</P><P class=""><STRONG>Our decision:</STRONG><SPAN>&nbsp;</SPAN>PSL is not a simple N:N connection.&nbsp; It carries<SPAN>&nbsp;</SPAN><STRONG>additional business information</STRONG><SPAN>&nbsp;</SPAN>(plant, validity period, priority), so it is<SPAN>&nbsp;</SPAN><STRONG>kept as a node</STRONG>. Two relationships are made explicit in the ontology.</P><PRE><CODE>bdc:fromSupplier a owl:ObjectProperty ; rdfs:domain bdc:PurchasingSourceList ; rdfs:range bdc:Supplier . bdc:forProduct a owl:ObjectProperty ; rdfs:domain bdc:PurchasingSourceList ; rdfs:range bdc:Product .</CODE></PRE><P class="">→ PSL is placed in the graph as the class<SPAN>&nbsp;</SPAN><CODE>bdc:PurchasingSourceList</CODE>, connected to both sides via<SPAN>&nbsp;</SPAN><CODE>bdc:fromSupplier</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>bdc:forProduct</CODE>.</P><P class="">→<SPAN>&nbsp;</SPAN><STRONG>The decision criterion is whether the junction represents a simple connection or carries additional meaning.</STRONG><SPAN>&nbsp;</SPAN>A junction that expresses only a simple connection should be flattened. The ontology result in Part 2-2, where PSL is classified as a<SPAN>&nbsp;</SPAN><CODE>bdc:StructuralNode</CODE><SPAN>&nbsp;</SPAN>(bridge node), reflects this decision.</P><H3 id="standard-mapping-vs-our-decisions-at-a-glance" id="toc-hId--768435712">Standard mapping vs our decisions: at a glance</H3><P>&nbsp;</P><TABLE><TBODY><TR><TD><STRONG>Standard default</STRONG></TD><TD><STRONG>Our domain decision</STRONG></TD><TD><STRONG>Where</STRONG></TD></TR><TR><TD>1 Table = 1 Class</TD><TD>Class Partition by discriminator column</TD><TD>Part 2-2</TD></TR><TR><TD>Reference triples only from direct FKs</TD><TD>Multi-hop precomputed into 1-hop via SQL Views</TD><TD>Part 3 Section 3</TD></TR><TR><TD>Junction table represented as an intermediate node</TD><TD>Meaningful junctions kept as nodes; simple connections flattened</TD><TD>Part 2-2</TD></TR></TBODY></TABLE><P class="">These three decisions are the<SPAN>&nbsp;</SPAN><STRONG>domain-level additions</STRONG><SPAN>&nbsp;</SPAN>we layered when translating the SAP BDC domain into a KG. The standard mapping provides a baseline; the domain decisions add business meaning on top.</P><P class="">The actual application of these decisions begins in the next post (Part 2-2).</P><HR /><H2 id="summary" id="toc-hId--671546210">Summary</H2><P class="">This post covered:</P><UL class=""><LI><STRONG>Core ontology concepts</STRONG>: Class/Instance (T-Box/A-Box), Domain &amp; Range, Competency Questions, and LLM affinity all rooted in Ontology 101.</LI><LI><STRONG>The seven-step process of Ontology 101</STRONG>: from determining the domain to creating instances. Iteration is the essence.</LI><LI><STRONG>W3C RDB→RDF mapping standards</STRONG>: Direct Mapping (automatic) and R2RML (manual and flexible). Table → Class, Row → Instance, Column → Literal Property, FK → ObjectProperty.</LI><LI><STRONG>A preview of domain decisions on top of the standards</STRONG>: Class Partition, multi-hop, junction handling, covered as concrete decisions in the next post.</LI></UL><HR /><H2 id="whats-next" id="toc-hId--868059715">What's Next</H2><P class="">The background is now in place. From here on, we move into the actual decisions we made for SAP BDC data.</P><P class=""><STRONG>Part 2-2 – Designing the Ontology for SAP Data</STRONG><SPAN>&nbsp;</SPAN>covers the domain decisions we made on top of these standards: which entities to place in the KG and which to leave in SQL, why we split a single Product table into three classes, what the final ontology structure looks like, and how we used AI to draft the ontology.</P><HR /><P class=""><EM>Stack: SAP BDC Data Products · HANA Cloud KGE · FastAPI · React · Claude (Anthropic / AI Core)</EM><SPAN>&nbsp;</SPAN><EM>Code:<SPAN>&nbsp;</SPAN><A href="https://github.com/claudiopark86/sap-bdc-kge-agent-workshop" target="_blank" rel="noopener nofollow noreferrer">github.com/claudiopark86/sap-bdc-kge-agent-workshop</A></EM></P><HR /><H2 id="references" id="toc-hId--896389529">References</H2><UL class=""><LI><STRONG>Ontology Development 101: A Guide to Creating Your First Ontology</STRONG><SPAN>&nbsp;</SPAN>by Noy &amp; McGuinness, Stanford 2001.<SPAN>&nbsp;</SPAN><A href="https://protege.stanford.edu/publications/ontology_development/ontology101.pdf" target="_blank" rel="noopener nofollow noreferrer">https://protege.stanford.edu/publications/ontology_development/ontology101.pdf</A></LI><LI><STRONG>A Direct Mapping of Relational Data to RDF</STRONG>, W3C Recommendation, 2012-09-27.<SPAN>&nbsp;</SPAN><A href="https://www.w3.org/TR/rdb-direct-mapping/" target="_blank" rel="noopener nofollow noreferrer">https://www.w3.org/TR/rdb-direct-mapping/</A></LI><LI><STRONG>R2RML: RDB to RDF Mapping Language</STRONG>, W3C Recommendation, 2012-09-27.<SPAN>&nbsp;</SPAN><A href="https://www.w3.org/TR/r2rml/" target="_blank" rel="noopener nofollow noreferrer">https://www.w3.org/TR/r2rml/</A></LI></UL> 2026-07-04T15:45:03.995000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-2-2-designing-the/ba-p/14274465 Knowledge Graph Agent on SAP HANA Cloud Series – Part 2-2: Designing the Ontology for SAP Data 2026-07-04T15:45:25.498000+02:00 ClaudioJP https://community.sap.com/t5/user/viewprofilepage/user-id/1509109 <P class="">This blog is part of a blog series on building AI Agents with SAP BDC (Business Data Cloud) Data Products and SAP HANA Cloud Knowledge Graph Engine:</P><UL><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-1-why-knowledge-graph/ba-p/14417381" target="_blank">Knowledge Graph Agent on SAP HANA Cloud Series – Part 1: Why Knowledge Graph for SAP Data</A></LI><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-2-1-ontology-standards/ba-p/14417385" target="_blank">Knowledge Graph Agent on SAP HANA Cloud Series – Part 2-1: Ontology Standards and Concepts</A></LI><LI><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/knowledge-graph-agent-on-sap-hana-cloud-series-part-2-2-designing-the/ba-p/14274465" target="_blank">Knowledge Graph Agent on SAP HANA Cloud Series – Part 2-2: Designing the Ontology for SAP Data</A>&nbsp;</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series – Part 3: Data Pipeline from BDC to HANA Cloud KGE</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series – Part 4: Building an Agent with SPARQL and SQL</LI><LI>Knowledge Graph Agent on SAP HANA Cloud Series&nbsp;– Part 5: Lessons Learned</LI></UL><P class=""><SPAN>Enterprise systems are vast, with thousands of tables, hundreds of business processes, and countless APIs, all interconnected in ways only domain experts fully grasp. For AI Agents to be genuinely useful in this environment, they need to understand not just what the data is, but how it all connects: which Product belongs to which Sales Order, which WorkCenter is located in which Plant, which Supplier serves which material. This understanding does not come from the data itself. It needs a dedicated semantic layer that makes business meaning explicit and machine-readable. Knowledge Graphs provide exactly that layer. SAP HANA Cloud now offers a native Knowledge Graph Engine to build it on top of SAP BDC Data Products. This series walks through the full journey: from why this layer matters, to how to design and build it, to how an AI Agent uses it to answer real business questions.</SPAN></P><P class="">This post covers the ontology design decisions we made for real SAP BDC data: which entities to place in the KG versus leave in SQL, why we split a single Product table into three classes, what the final ontology structure looks like, and how we used AI to draft the Turtle file, including multilingual labels, comments, and synonyms.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span><SPAN>&nbsp;</SPAN><STRONG>Reading guide.</STRONG><SPAN>&nbsp;</SPAN>This post is designed to stand on its own. Section 0 below briefly covers the minimum KG modeling vocabulary required to follow along (the table/row/column/relationship-link mapping and why<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>matters). Section 1 prepares the environment, and the substantive work begins in Section 2. If you want the deeper W3C standards background and the conceptual foundation, consult<SPAN>&nbsp;</SPAN><STRONG>Part 2-1 – Ontology Standards and Concepts</STRONG>.</P></BLOCKQUOTE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":file_folder:">📁</span><SPAN>&nbsp;</SPAN><STRONG>Code reference.</STRONG><SPAN>&nbsp;</SPAN>All scripts and project files mentioned in this post are available in the demo repository:<SPAN>&nbsp;</SPAN><A href="https://github.com/claudiopark86/sap-bdc-kge-agent-workshop" target="_blank" rel="noopener nofollow noreferrer">github.com/claudiopark86/sap-bdc-kge-agent-workshop</A></P><P class="">File names such as<SPAN>&nbsp;</SPAN><CODE>00_hana_setup.sql</CODE>,<SPAN>&nbsp;</SPAN><CODE>01_setup_synonyms.py</CODE>,<SPAN>&nbsp;</SPAN><CODE>02_verify_joins.py</CODE>,<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE>, and<SPAN>&nbsp;</SPAN><CODE>04_generate_ontology.py</CODE><SPAN>&nbsp;</SPAN>refer to files in that repository.</P></BLOCKQUOTE><P class=""><STRONG>Table of contents:</STRONG></P><OL class=""><LI>KG vocabulary essentials</LI><LI>Preparing the work environment: BDC connection, users, synonyms</LI><LI>Organizing your data as ontology material</LI><LI>Deciding between KG and SQL</LI><LI>Splitting one Data Product into multiple classes (Class Partition)</LI><LI>Completing the ontology</LI><LI>Drafting the ontology Turtle with AI, and writing labels</LI></OL><HR /><H1 id="0-kg-vocabulary-essentials" id="toc-hId-1636489576">0. KG Vocabulary Essentials</H1><P class="">Before moving into the work itself, here is the minimum KG modeling vocabulary required to follow this post. If you have read Part 2-1 in detail, this is a brief refresher; if you are seeing it for the first time, this much is sufficient to follow Sections 2–6.</P><H3 id="01-relational-db--kg-mapping-at-a-glance" id="toc-hId-1698141509">0.1 Relational DB → KG mapping (at a glance)</H3><P class="">The core mapping between a relational database and an RDF Knowledge Graph, as defined by the W3C:</P><TABLE><TBODY><TR><TD><STRONG>Relational DB</STRONG></TD><TD><STRONG>Knowledge Graph</STRONG></TD><TD><STRONG>Example</STRONG></TD></TR><TR><TD><STRONG>Table</STRONG></TD><TD><STRONG>Class</STRONG><SPAN>&nbsp;</SPAN>(concept)</TD><TD><CODE>WORK_CENTER</CODE><SPAN>&nbsp;</SPAN>table →<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter</CODE><SPAN>&nbsp;</SPAN>class</TD></TR><TR><TD><STRONG>Row</STRONG></TD><TD><STRONG>Instance</STRONG><SPAN>&nbsp;</SPAN>(actual entity)</TD><TD>Row Z_ASM3 →<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter_Z_ASM3</CODE><SPAN>&nbsp;</SPAN>instance</TD></TR><TR><TD><STRONG>Column</STRONG></TD><TD><STRONG>DatatypeProperty</STRONG><SPAN>&nbsp;</SPAN>(attribute / literal property)</TD><TD><CODE>WorkCenterName</CODE><SPAN>&nbsp;</SPAN>→<SPAN>&nbsp;</SPAN><CODE>bdc:workCenterName "Z_ASM3"</CODE></TD></TR><TR><TD><STRONG>Relationship link (FK and similar)</STRONG></TD><TD><STRONG>ObjectProperty</STRONG><SPAN>&nbsp;</SPAN>(relationship)</TD><TD><CODE>PlantID</CODE><SPAN>&nbsp;</SPAN>→<SPAN>&nbsp;</SPAN><CODE>bdc:locatedAt bdc:Plant_1010</CODE></TD></TR></TBODY></TABLE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":warning:">⚠️</span><SPAN>&nbsp;</SPAN><STRONG>This table shows the basic mapping principles.</STRONG><SPAN>&nbsp;</SPAN>In practice, the correspondence is rarely 1:1. A single table may be split into multiple classes (for example, splitting<SPAN>&nbsp;</SPAN><CODE>MaterialType</CODE><SPAN>&nbsp;</SPAN>into FinishedGood/SemiFinished/RawMaterial via Class Partition), or conversely, multiple tables may form a single class (for example, WorkCenter + WorkCenterCapacity →<SPAN>&nbsp;</SPAN><CODE>bdc:WorkCenter</CODE>). These domain decisions are covered in<SPAN>&nbsp;</SPAN><STRONG>Section 4 of this Part</STRONG><SPAN>&nbsp;</SPAN>and in<SPAN>&nbsp;</SPAN><STRONG>Part 2-1 Sections 3–4</STRONG>.</P></BLOCKQUOTE><P class="">The key point is that<SPAN>&nbsp;</SPAN><STRONG>relationship links are what make a graph a graph</STRONG>. More precisely: an FK (or a column shared between two tables) is a<SPAN>&nbsp;</SPAN><STRONG>hint</STRONG><SPAN>&nbsp;</SPAN>that<SPAN>&nbsp;</SPAN><EM>"these two entities can be connected."</EM><SPAN>&nbsp;</SPAN>Which business meaning to assign to that connection, and which name (ObjectProperty) to use, is a<SPAN>&nbsp;</SPAN><STRONG>decision made by a domain expert</STRONG>. Direct Mapping looks at an FK and automatically produces a reference triple of the form<SPAN>&nbsp;</SPAN><CODE>#ref-colName</CODE>; in R2RML, a person explicitly defines a business term such as<SPAN>&nbsp;</SPAN><CODE>bdc:locatedAt</CODE>. The details are covered in<SPAN>&nbsp;</SPAN><STRONG>Part 2-1 Section 3</STRONG>.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":open_book:">📖</span><SPAN>&nbsp;</SPAN><STRONG>Standard basis:</STRONG><SPAN>&nbsp;</SPAN>The mapping above is grounded in the W3C<SPAN>&nbsp;</SPAN><EM>Direct Mapping</EM><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><EM>R2RML</EM><SPAN>&nbsp;</SPAN>recommendations (2012). The full standard definitions and the standard's limitations are covered in<SPAN>&nbsp;</SPAN><STRONG>Part 2-1 Sections 3–4</STRONG>.</P></BLOCKQUOTE><H3 id="02-rdfslabel-is-the-minimum-requirement-for-a-kg" id="toc-hId-1501628004">0.2<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>is the minimum requirement for a KG</H3><P class="">Once instances are loaded into a KG, SPARQL results return URIs:</P><PRE><CODE>bdc:Supplier_1000015 bdc:Product_MZ-FG-C900 bdc:WorkCenter_Z_ASM3</CODE></PRE><P class="">An LLM given only URIs has a hard time interpreting<SPAN>&nbsp;</SPAN><EM>"what this is."</EM><SPAN>&nbsp;</SPAN><STRONG>Attach<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>to every instance, class, and property</STRONG>, and the SPARQL results return human-readable names directly.</P><PRE><CODE>wcLabel | productLabel Z_ASM3 | C-Series Road Bike 900 Z_ASM3 | C-Series Road Bike 950</CODE></PRE><P class=""><STRONG>Without labels, the KG functions, but the Agent cannot interpret the results.</STRONG><SPAN>&nbsp;</SPAN>Labels are the minimum requirement that makes a KG into<SPAN>&nbsp;</SPAN><EM>"a graph an LLM can read."</EM></P><H4 id="not-only-labels-but-comments-and-not-only-on-instances-but-on-classes-and-relationships-too" id="toc-hId-1434197218">Not only labels but comments, and not only on instances but on classes and relationships too</H4><P class="">The instance example above illustrates<SPAN>&nbsp;</SPAN><EM>"why a label is needed."</EM><SPAN>&nbsp;</SPAN>In an actual ontology, one more property is required, and the places where label and comment are attached extend further.</P><UL class=""><LI><STRONG><CODE>rdfs:comment</CODE></STRONG>: where label provides<SPAN>&nbsp;</SPAN><EM>"a short name,"</EM><SPAN>&nbsp;</SPAN>comment is a natural-language description. A one-line answer to<SPAN>&nbsp;</SPAN><EM>"what is a WorkCenter?"</EM></LI><LI><STRONG>Attached to classes and relationships (ObjectProperty) as well</STRONG>: not only on instances but on<SPAN>&nbsp;</SPAN><STRONG>the class definitions themselves</STRONG><SPAN>&nbsp;</SPAN>(T-Box). Tools such as Protégé and Metaphactory display labels instead of URIs, and when the Part 4 Agent loads ontology metadata into its LLM context, comments serve as the explanation of<SPAN>&nbsp;</SPAN><EM>"what role this class plays."</EM></LI></UL><PRE><CODE>bdc:WorkCenter a owl:Class ; rdfs:label "WorkCenter"@en, "작업장"@ko ; rdfs:comment "Production work center. Located in a Plant, produces Finished Goods."@en, "생산 작업장. 특정 Plant 에 위치하며 FinishedGood 을 생산한다."@ko . bdc:producesProduct a owl:ObjectProperty ; rdfs:label "produces product"@en, "제품 생산"@ko ; rdfs:comment "The FinishedGood that a WorkCenter produces."@en ; rdfs:domain bdc:WorkCenter ; rdfs:range bdc:FinishedGood .</CODE></PRE><BLOCKQUOTE dir="auto"><P class=""><SPAN>&nbsp;</SPAN><STRONG>Standard basis:</STRONG>&nbsp;OWL 2 Quick Reference §2.7 defines <CODE>rdfs:label</CODE> and <CODE>rdfs:comment</CODE> as built-in annotation properties. These are "standard annotations that can be attached to any class, property, or instance." Major ontologies such as FOAF, schema.org, and Dublin Core commonly attach them to T-Box definitions.</P></BLOCKQUOTE><P class="">The practical convention differs slightly depending on where they are attached:</P><TABLE><TBODY><TR><TD><STRONG>Where</STRONG></TD><TD><STRONG>label</STRONG></TD><TD><STRONG>comment</STRONG></TD></TR><TR><TD>Class · relationship (ObjectProperty) (T-Box)</TD><TD>Nearly always — business name, verb phrase</TD><TD>Nearly always — concept description, LLM context</TD></TR><TR><TD>Instance (A-Box)</TD><TD>Nearly always — identifier or name (<CODE>"Z_ASM3"</CODE>)</TD><TD>Rarely — typically replaced by a dedicated datatype property</TD></TR></TBODY></TABLE><P class="">Comments on the A-Box are rare because facts about an instance are more useful as<SPAN>&nbsp;</SPAN><STRONG>discrete datatype properties</STRONG><SPAN>&nbsp;</SPAN>than as free-form prose. They are easier to validate and query. Rather than the free-form<SPAN>&nbsp;</SPAN><EM>"Z_ASM3 is the bicycle final-assembly line,"</EM><SPAN>&nbsp;</SPAN>the same information is split into<SPAN>&nbsp;</SPAN><CODE>bdc:workCenterTypeCode "A"</CODE><SPAN>&nbsp;</SPAN>or similar.</P><H4 id="internationalization-i18n-language-tags-in-a-single-property" id="toc-hId-1237683713">Internationalization (i18n): language tags in a single property</H4><P class=""><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>rdfs:comment</CODE><SPAN>&nbsp;</SPAN>accept RDF's<SPAN>&nbsp;</SPAN><STRONG>language-tagged literals</STRONG><SPAN>&nbsp;</SPAN>directly. Multiple language values can sit on the same property with<SPAN>&nbsp;</SPAN><CODE>@en</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>@ko</CODE><SPAN>&nbsp;</SPAN>and other language tags (as in the Turtle example above). With this in place, the Part 4 Agent retrieves the user's language with a single line <SPAN>&nbsp;</SPAN><CODE>FILTER(LANG(?label) = "ko")</CODE><SPAN>&nbsp;</SPAN>&nbsp;in SPARQL. Multilingual support is delivered through the ontology alone, with no backend code changes.</P><H4 id="when-users-call-the-same-entity-by-different-names-skosaltlabel" id="toc-hId-1041170208">When users call the same entity by different names: skos:altLabel</H4><P class="">One more situation comes up:<SPAN>&nbsp;</SPAN><STRONG>when users refer to the same entity by different names</STRONG>. Calling a<SPAN>&nbsp;</SPAN><EM>"공급사"</EM><SPAN>&nbsp;</SPAN>a<SPAN>&nbsp;</SPAN><EM>"벤더,"</EM><SPAN>&nbsp;</SPAN>or using<SPAN>&nbsp;</SPAN><EM>"Vendor"</EM><SPAN>&nbsp;</SPAN>instead of<SPAN>&nbsp;</SPAN><EM>"Supplier."</EM><SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>is a vocabulary for assigning one preferred name per entity, so packing synonyms into it is awkward. This is where the W3C<SPAN>&nbsp;</SPAN><STRONG>SKOS</STRONG><SPAN>&nbsp;</SPAN>(Simple Knowledge Organization System) standard's<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>fills the gap.</P><PRE><CODE>bdc:Supplier a owl:Class ; rdfs:label "Supplier"@en, "공급사"@ko ; skos:altLabel "Vendor"@en, "벤더"@ko, "거래처"@ko .</CODE></PRE><P class="">When the Agent encounters an expression such as<SPAN>&nbsp;</SPAN><EM>"vendor USSU-VSF01..."</EM><SPAN>&nbsp;</SPAN>in a user question, the<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>in the ontology links it to<SPAN>&nbsp;</SPAN><CODE>bdc:Supplier</CODE>.<SPAN>&nbsp;</SPAN><STRONG>If<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>is the "official name,"<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>holds "the other recognized names."</STRONG><SPAN>&nbsp;</SPAN>The SAP domain mixes ABAP abbreviations (<CODE>EKKO</CODE>,<SPAN>&nbsp;</SPAN><CODE>KNA1</CODE>) with business vocabulary (<CODE>PurchaseOrder</CODE>,<SPAN>&nbsp;</SPAN><CODE>Customer</CODE>), so this pattern fits especially well.</P><P class="">How labels, comments, and altLabels are applied in bulk is covered in Section 6. Rather than typing each one by hand for dozens of classes, an LLM given the YAML and domain context fills them in consistently.</P><H1 id="1-preparing-the-work-environment-bdc-connection-users-synonyms" id="toc-hId-457408546">1. Preparing the Work Environment: BDC Connection, Users, Synonyms</H1><P class="">With the vocabulary established, we can now prepare the implementation environment. Before substantive ontology work begins, the environment must be ready. Three pieces are required:</P><OL class=""><LI><STRONG>Source:</STRONG><SPAN>&nbsp;</SPAN>Install an SAP BDC Data Product into HANA Cloud to obtain the Virtual Table and CSN metadata.</LI><LI><STRONG>User &amp; Permissions:</STRONG><SPAN>&nbsp;</SPAN>A working user with catalog read, Virtual Table SELECT, and ownership of the working schema.</LI><LI><STRONG>Synonym:</STRONG><SPAN>&nbsp;</SPAN>A mapping that groups long Virtual Table names into short ones to make subsequent SQL queries and KG generation work easier.</LI></OL><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":information:">ℹ️</span><SPAN>&nbsp;</SPAN><STRONG>About the<SPAN>&nbsp;</SPAN><CODE>.env</CODE><SPAN>&nbsp;</SPAN>file</STRONG></P><P class="">The<SPAN>&nbsp;</SPAN><CODE>.env</CODE><SPAN>&nbsp;</SPAN>referenced throughout this post is a<SPAN>&nbsp;</SPAN><STRONG>local-development environment variable file</STRONG>. It holds sensitive values, including HANA Cloud connection details (<CODE>HANA_ADDRESS</CODE>,<SPAN>&nbsp;</SPAN><CODE>HANA_PORT</CODE>,<SPAN>&nbsp;</SPAN><CODE>HANA_USER</CODE>,<SPAN>&nbsp;</SPAN><CODE>HANA_PASSWORD</CODE>), SAP AI Core keys, AWS S3 keys, which the LLM references while writing code.<SPAN>&nbsp;</SPAN><STRONG>It is required only for local testing</STRONG>; for production deployments, use a secret management service such as Kyma Secret.</P><P class="">All Vibe Coding prompts in this post assume<SPAN>&nbsp;</SPAN><CODE>.env</CODE><SPAN>&nbsp;</SPAN>is already configured.</P></BLOCKQUOTE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":information:">ℹ️</span><SPAN>&nbsp;</SPAN><STRONG>About<SPAN>&nbsp;</SPAN><CODE>utils.py</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>get_hana_connection()</CODE></STRONG></P><P class="">The<SPAN>&nbsp;</SPAN><CODE>utils.py</CODE><SPAN>&nbsp;</SPAN>referenced in the Vibe Coding prompts is a<SPAN>&nbsp;</SPAN><STRONG>helper module that centralizes connections to external systems</STRONG><SPAN>&nbsp;such as&nbsp;</SPAN>HANA Cloud, AI Core, S3.<SPAN>&nbsp;</SPAN><CODE>get_hana_connection()</CODE><SPAN>&nbsp;</SPAN>is the function that builds a HANA Cloud connection: it reads the<SPAN>&nbsp;</SPAN><CODE>.env</CODE><SPAN>&nbsp;</SPAN>connection details and returns an<SPAN>&nbsp;</SPAN><CODE>hdbcli</CODE><SPAN>&nbsp;</SPAN>connection object.</P><P class="">Because the LLM can call this single function instead of rewriting connection code each time, the resulting code is shorter and the only thing that needs to change when the environment shifts is<SPAN>&nbsp;</SPAN><CODE>.env</CODE>. Copy our demo's<SPAN>&nbsp;</SPAN><CODE>utils.py</CODE><SPAN>&nbsp;</SPAN>into your project, or ask the LLM directly:<SPAN>&nbsp;</SPAN><EM>"Create a<SPAN>&nbsp;</SPAN><CODE>utils.py</CODE><SPAN>&nbsp;</SPAN>with a<SPAN>&nbsp;</SPAN><CODE>get_hana_connection()</CODE><SPAN>&nbsp;</SPAN>function that reads<SPAN>&nbsp;</SPAN><CODE>.env</CODE><SPAN>&nbsp;</SPAN>and returns a HANA connection."</EM></P></BLOCKQUOTE><H3 id="source-connecting-a-bdc-data-product-to-hana-cloud" id="toc-hId-519060479">Source: connecting a BDC Data Product to HANA Cloud</H3><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":pushpin:">📌</span><SPAN>&nbsp;</SPAN><STRONG>This section covers the hands-on setup that Part 1 Section6 introduced.</STRONG><SPAN>&nbsp;</SPAN></P><P class=""><SPAN>&nbsp;</SPAN><STRONG>References.</STRONG><SPAN>&nbsp;</SPAN>This series does not cover the detailed procedure from installing BDC Data Products into HANA Cloud and getting CSNs of them from SAP HANA Cloud Central as Part 1 already explained them. Please, check out the Part 1 and the the official SAP guides.</P><UL class=""><LI>Tutorial:<SPAN>&nbsp;</SPAN><A href="https://developers.sap.com/tutorials/hana-cloud-data-products-consumption.html" target="_blank" rel="noopener noreferrer">HANA Cloud Data Products Consumption</A></LI><LI>Official documentation:<SPAN>&nbsp;</SPAN><A href="https://help.sap.com/docs/hana-cloud/sap-hana-cloud-administration-guide/data-product-support-in-sap-hana-cloud-internal?locale=en-US" target="_blank" rel="noopener noreferrer">Data Product Support in SAP HANA Cloud</A></LI></UL><P>Here, you will find the steps: after download the CSN JSON files, where to put them and how to set up the HANA Cloud environment before ontology work begins.</P></BLOCKQUOTE><P class="">Ontology design begins the moment an SAP BDC Data Product is shared into HANA Cloud. Select a Data Product in the BDC catalog and install it into HANA Cloud: a Virtual Table is created, and at the same time the<SPAN>&nbsp;</SPAN><STRONG>CSN metadata</STRONG><SPAN>&nbsp;</SPAN>(column names, types, FK associations, and so on) becomes available. With no separate script for extracting metadata, the raw material for ontology design is in place after a few clicks.</P><P class="">This metadata is the starting point for ontology design. When you write an ontology, you need to know precisely which entities, columns, and relationships exist, and which Competency Questions (business requirements) the ontology must address. The metadata addresses the first half.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":file_folder:">📁</span><SPAN>&nbsp;</SPAN><STRONG>CSN metadata is gathered in the<SPAN>&nbsp;</SPAN><CODE>docs/dp_specs/</CODE><SPAN>&nbsp;</SPAN>folder.</STRONG></P><P class="">When you install each Data Product from BDC into HANA Cloud, you can download the CSN JSON file alongside it (from the BDC Catalog Data Product detail view, or the HANA Cloud Central Data Product tab). Place the downloaded files in your project's<SPAN>&nbsp;</SPAN><CODE>docs/dp_specs/</CODE><SPAN>&nbsp;</SPAN>folder, and the LLM can reference them in bulk during the later Vibe Coding stages as<SPAN>&nbsp;</SPAN><CODE>docs/dp_specs/*_CSN.json</CODE>.</P><PRE><CODE>docs/dp_specs/ ├── WorkCenter_CSN.json ├── Product_CSN.json ├── Supplier_CSN.json ├── Customer_CSN.json └── ...</CODE></PRE><P class="">This folder becomes the shared input for Section 2 (CSN_SUMMARY authoring), Section 3 (KG / SQL classification), and Section 4 (Class Partition decisions).</P></BLOCKQUOTE><P class="">For teams working directly with S/4HANA, as covered in Part 1 Section 1, CDS View definitions can serve as the same metadata raw material.</P><H3 id="user--permissions-a-working-user-with-catalog-read-access" id="toc-hId-322546974">User &amp; Permissions: a working user with catalog read access</H3><P class="">To query installed Virtual Tables and use them for ontology work, a user with appropriate permissions is required. Rather than performing every operation as DBADMIN, it is cleaner to create a dedicated working user (for example,<SPAN>&nbsp;</SPAN><CODE>BDC_DP</CODE>) and grant the following:</P><UL class=""><LI>Read access to metadata views such as<SPAN>&nbsp;</SPAN><CODE>SYS.VIRTUAL_TABLES</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>PUBLIC.VIRTUAL_COLUMNS</CODE>: for Virtual Table and column metadata exploration</LI><LI><CODE>SELECT</CODE><SPAN>&nbsp;</SPAN>on each BDC DP schema: for Virtual Table queries</LI><LI>Ownership of a working schema (for example,<SPAN>&nbsp;</SPAN><CODE>BDC_DP</CODE><span class="lia-unicode-emoji" title=":disappointed_face:">😞</span> where synonyms and KG helper Views will be created</LI></UL><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span>&nbsp;Our demo's<SPAN>&nbsp;</SPAN><CODE>00_hana_setup.sql</CODE><SPAN>&nbsp;</SPAN>performs the same setup in one shot. It bundles user creation, BDC DP schema SELECT grants, system catalog read access, and working-schema creation into a single SQL script.</P></BLOCKQUOTE><H3 id="synonym-shortening-long-virtual-table-names" id="toc-hId-126033469">Synonym: shortening long Virtual Table names</H3><BLOCKQUOTE dir="auto"><P class="">This step is genuinely needed only after Section 3, once you have decided<SPAN>&nbsp;</SPAN><EM>"which Data Products and which tables within them to place in the KG."</EM><SPAN>&nbsp;</SPAN>It is sufficient to create synonyms only for the chosen tables. Still, because the flow<SPAN>&nbsp;</SPAN><EM>"BDC connection → user setup → make SQL easier to write"</EM><SPAN>&nbsp;</SPAN>reads naturally as one block, and because subsequent work flows more smoothly afterward, we cover it here in advance. Whether to define synonyms only for selected tables or for every Virtual Table brought from BDC is a matter of working style.</P></BLOCKQUOTE><P class="">Immediately after installation, two names are auto-generated.<SPAN>&nbsp;</SPAN><STRONG>A separate schema is created per Data Product</STRONG>, and the Virtual Table sits inside it:</P><PRE><CODE>DP schema name: _SAP_DATAPRODUCT_sap_s4com_dataProduct_WorkCenter_v1_&lt;DP_GUID&gt; Virtual Table inside it: _SAP_DATAPRODUCT_&lt;VT_GUID&gt;_workcenter.WorkCenter</CODE></PRE><P class="">The schema itself is long, multiple DPs produce many schemas, and the table names inside each schema are long as well. Writing these out in every SQL statement is inefficient.<SPAN>&nbsp;</SPAN><STRONG>Consolidating them into short synonyms in one place</STRONG><SPAN>&nbsp;</SPAN>is the first step before substantive work begins.</P><PRE><CODE><SPAN class="">CREATE</SPAN> SYNONYM BDC_DP.WORK_CENTER <SPAN class="">FOR</SPAN> "_SAP_DATAPRODUCT_&lt;VT_GUID&gt;_workcenter"."WorkCenter";</CODE></PRE><P class="">With this in place, subsequent SQL can use the short name<SPAN>&nbsp;</SPAN><CODE>BDC_DP.WORK_CENTER</CODE>.</P><H4 id="vibe-coding-querying-the-virtual-table-list" id="toc-hId--439114412">Vibe Coding: querying the Virtual Table list</H4><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>.env</CODE>: HANA Cloud connection details (HANA_ADDRESS, HANA_PORT, HANA_USER, HANA_PASSWORD)</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI><CODE>virtual_tables.csv</CODE>: full list of available BDC Virtual Tables (SCHEMA_NAME, TABLE_NAME)</LI><LI>Used in: Section 1 synonym creation, Section 2 CSN_SUMMARY authoring</LI></UL></BLOCKQUOTE><P class="">This step is delegated to the LLM. The human reviewer's role is limited to confirming that the resulting CSV is produced correctly.</P><PRE><CODE>Make get_hana_connection() in utils.py read the HANA Cloud connection details (HANA_ADDRESS, HANA_PORT, HANA_USER, HANA_PASSWORD) from .env and connect to HANA Cloud. Then use this function to query SCHEMA_NAME and TABLE_NAME from the SYS.VIRTUAL_TABLES view, and save the result as virtual_tables.csv.<BR /><BR />Use docs/dp_specs/virtual_tables_template.csv as the column structure reference (SCHEMA_NAME, TABLE_NAME — synonym columns will be added in the next step)</CODE></PRE><H4 id="vibe-coding-synonym-creation" id="toc-hId--635627917">Vibe Coding: synonym creation</H4><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>virtual_tables.csv</CODE>: the VT list from the previous step (SCHEMA_NAME, TABLE_NAME)</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI><CODE>ingestion/01_setup_synonyms.py</CODE>: synonym creation script</LI><LI>Short-name synonyms created in HANA's<SPAN>&nbsp;</SPAN><CODE>BDC_DP</CODE><SPAN>&nbsp;</SPAN>schema (long SAP table names → short names)</LI><LI><CODE>virtual_tables.csv</CODE><SPAN>&nbsp;</SPAN>updated: enriched CSV with<SPAN>&nbsp;</SPAN><CODE>synonym_schema</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>synonym_name</CODE><SPAN>&nbsp;</SPAN>columns added</LI><LI>Used in: Section 2 CSN_SUMMARY authoring, Section 2 JOIN verification, the full Part 3 pipeline</LI></UL></BLOCKQUOTE><P class="">The LLM examines the VT list, proposes short synonym names, and generates the code for creating them. The human reviewer confirms that there are no collisions (no existing synonym with the same name) and that the short names match the domain vocabulary, then runs the script in the local development environment. Once the names are recorded in the CSV, subsequent stages can carry the VT ↔ synonym mapping in a single file.</P><PRE><CODE>For each VT in virtual_tables.csv, propose a short synonym name (for example, WORK_CENTER, SALES_ORDER_ITEM) and write the code that creates each synonym under my working schema (BDC_DP) into ingestion/01_setup_synonyms.py. Run it immediately after writing, and update virtual_tables.csv by adding synonym_schema and synonym_name columns.<BR />The final CSV should match the 4-column structure in docs/dp_specs/virtual_tables_template.csv (SCHEMA_NAME, TABLE_NAME, synonym_schema, synonym_name).</CODE></PRE><H3 id="section-1-deliverables" id="toc-hId--538738415">Section 1 deliverables</H3><TABLE><TBODY><TR><TD width="154.641px" height="57px"><STRONG>Step</STRONG></TD><TD width="370.336px" height="57px"><STRONG>Deliverable</STRONG></TD><TD width="360.023px" height="57px"><STRONG>Role</STRONG></TD></TR><TR><TD width="154.641px" height="57px">BDC → HANA install (manual)</TD><TD width="370.336px" height="57px"><CODE>docs/dp_specs/*_CSN.json</CODE></TD><TD width="360.023px" height="57px">Input for Section 2 CSN_SUMMARY</TD></TR><TR><TD width="154.641px" height="57px">User &amp; Permissions (manual SQL)</TD><TD width="370.336px" height="57px">Working user + working schema</TD><TD width="360.023px" height="57px">Context for all subsequent SQL / synonym creation</TD></TR><TR><TD width="154.641px" height="57px">Query VT list (Vibe Coding)</TD><TD width="370.336px" height="57px"><CODE>virtual_tables.csv</CODE></TD><TD width="360.023px" height="57px">Input for synonym creation and CSN_SUMMARY VT matching</TD></TR><TR><TD width="154.641px" height="85px">Create synonyms (Vibe Coding)</TD><TD width="370.336px" height="85px"><CODE>01_setup_synonyms.py</CODE><SPAN>&nbsp;</SPAN>+ HANA synonyms + enriched<SPAN>&nbsp;</SPAN><CODE>virtual_tables.csv</CODE></TD><TD width="360.023px" height="85px">Short names for SQL queries and later KG generation steps, the working environment for Sections 2–6</TD></TR></TBODY></TABLE><P class="">Section 1 ends here as<SPAN>&nbsp;</SPAN><STRONG>environment preparation</STRONG>. The substantive question,<SPAN>&nbsp;</SPAN><EM>"what to design on top of this data"</EM><SPAN>&nbsp;</SPAN>(KG vs SQL classification, classes, relationships), begins at Section 2 with the CSN_SUMMARY and accumulates in<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE>.</P><H1 id="2-organizing-your-data-as-ontology-material" id="toc-hId--148445906">2. Organizing Your Data as Ontology Material</H1><P class="">With the environment ready, the actual ontology work begins. Authoring the CSN_SUMMARY, drafting<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE>, and verifying JOIN candidates for relationships fall within the scope of Section 2.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":direct_hit:">🎯</span><SPAN>&nbsp;</SPAN><STRONG>What you will build starting in Section 2:<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>(the single source of truth for ontology work)</STRONG></P><P class="">If Section 1 was environment setup, Section 2 starts the<SPAN>&nbsp;</SPAN><STRONG>design decisions</STRONG>. The file in which those decisions accumulate is<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE>, and it is the<SPAN>&nbsp;</SPAN><STRONG>single source of truth that collects every decision made across Sections 2–6</STRONG>.<SPAN>&nbsp;</SPAN><EM>"Which DPs go into the KG, which entities become which classes, how to partition".&nbsp;All of these decisions converge in one file.</EM>&nbsp;The blank template is in<SPAN>&nbsp;</SPAN><CODE>ingestion/dp_mapping.template.yaml</CODE>, and the final form follows a five-region structure:</P><PRE><CODE><SPAN class=""># ingestion/dp_mapping.template.yaml (excerpt, summarized)</SPAN> <SPAN class="">synonym_schema:</SPAN> <SPAN class="">&amp;synonym_schema</SPAN> <SPAN class="">"BDC_DP"</SPAN> <SPAN class=""># the synonym schema created in Section 1</SPAN> <SPAN class=""># ─── Master Data: core node DPs that serve as endpoints of relationships (the start/end of CQs) ───</SPAN> <SPAN class="">master_data:</SPAN> <SPAN class="">-</SPAN> <SPAN class="">name:</SPAN> <SPAN class="">&lt;DP.</SPAN> <SPAN class="">e.g.,</SPAN> <SPAN class="">WorkCenter&gt;</SPAN> <SPAN class="">kg_class:</SPAN> <SPAN class="">&lt;bdc:ClassName&gt;</SPAN> <SPAN class="">description:</SPAN> <SPAN class="">"&lt;one-line DP description&gt;"</SPAN> <SPAN class="">virtual_table_schema:</SPAN> <SPAN class="">"_SAP_DATAPRODUCT_..._&lt;DP&gt;_v1_&lt;DP_GUID&gt;"</SPAN> <SPAN class="">entities:</SPAN> <SPAN class="">-</SPAN> <SPAN class="">entity:</SPAN> <SPAN class="">&lt;entity</SPAN> <SPAN class="">name&gt;</SPAN> <SPAN class="">synonym_schema:</SPAN> <SPAN class="">*synonym_schema</SPAN> <SPAN class="">synonym_name:</SPAN> <SPAN class="">"&lt;SHORT_NAME&gt;"</SPAN> <SPAN class="">virtual_table_name:</SPAN> <SPAN class="">"_SAP_DATAPRODUCT_&lt;VT_GUID&gt;_&lt;dp&gt;.&lt;Entity&gt;"</SPAN> <SPAN class=""># partition [Optional]: added for DPs that require Class Partition (Section 4 decision)</SPAN> <SPAN class="">partition:</SPAN> <SPAN class="">discriminator:</SPAN> <SPAN class="">&lt;column&gt;</SPAN> <SPAN class="">parent_class:</SPAN> <SPAN class="">&lt;kg_class</SPAN> <SPAN class="">above&gt;</SPAN> <SPAN class="">subclasses:</SPAN> <SPAN class="">-</SPAN> <SPAN class="">value:</SPAN> <SPAN class="">&lt;column</SPAN> <SPAN class="">value&gt;</SPAN> <SPAN class="">kg_class:</SPAN> <SPAN class="">&lt;bdc:SubClass&gt;</SPAN> <SPAN class=""># ─── Structural Data [Optional]: multi-hop bridge / junction roles ───</SPAN> <SPAN class=""># If master_data alone connects all relationships, this category can stay empty.</SPAN> <SPAN class="">structural_data:</SPAN> <SPAN class="">-</SPAN> <SPAN class="">name:</SPAN> <SPAN class="">&lt;DP&gt;</SPAN> <SPAN class="">role:</SPAN> <SPAN class="">&lt;multi-hop-intermediate</SPAN> <SPAN class="">|</SPAN> <SPAN class="">bridge-with-attributes&gt;</SPAN> <SPAN class="">...</SPAN> <SPAN class=""># ─── Transaction Data: not loaded into the KG, used as SQL aggregation targets ───</SPAN> <SPAN class="">transaction_data:</SPAN> <SPAN class="">-</SPAN> <SPAN class="">name:</SPAN> <SPAN class="">&lt;DP&gt;</SPAN> <SPAN class="">aggregation_columns:</SPAN> [<SPAN class="">&lt;col1&gt;</SPAN>, <SPAN class="">&lt;col2&gt;</SPAN>] <SPAN class="">...</SPAN> <SPAN class=""># ─── Relationships: ObjectProperty definitions + KGR View mapping (Section 5 decision) ───</SPAN> <SPAN class="">relationships:</SPAN> <SPAN class="">-</SPAN> <SPAN class="">name:</SPAN> <SPAN class="">&lt;relationship</SPAN> <SPAN class="">name,</SPAN> <SPAN class="">e.g.,</SPAN> <SPAN class="">producesProduct&gt;</SPAN> <SPAN class="">domain:</SPAN> <SPAN class="">&lt;bdc:Class&gt;</SPAN> <SPAN class="">range:</SPAN> <SPAN class="">&lt;bdc:Class&gt;</SPAN> <SPAN class="">source:</SPAN> <SPAN class="">type:</SPAN> <SPAN class="">kgr_view</SPAN> <SPAN class="">view_name:</SPAN> <SPAN class="">&lt;KGR_VIEW_NAME&gt;</SPAN> <SPAN class="">sql_file:</SPAN> <SPAN class="">&lt;path,</SPAN> <SPAN class="">e.g.,</SPAN> <SPAN class="">ingestion/kgr_views/kgr_*.sql&gt;</SPAN> </CODE></PRE><P class="">Because this YAML collects every decision in Sections 2–6, it appears repeatedly in the rest of this post. Following<SPAN>&nbsp;</SPAN><EM>"where in the YAML each step's decision is recorded"</EM><SPAN>&nbsp;</SPAN>gives a natural view of how the ontology comes together. For the full template and field descriptions, consult<SPAN>&nbsp;</SPAN><CODE>ingestion/dp_mapping.template.yaml</CODE><SPAN>&nbsp;</SPAN>directly.</P><P class="">One clarification:<SPAN>&nbsp;</SPAN><STRONG>the YAML does not hold everything.</STRONG><SPAN>&nbsp;</SPAN>The Turtle definitions of classes and properties themselves (including label and comment) are emitted in Section 6 as a separate<SPAN>&nbsp;</SPAN><CODE>ontology.ttl</CODE>, and the SQL bodies of the KGR Views live in separate<SPAN>&nbsp;</SPAN><CODE>.sql</CODE><SPAN>&nbsp;</SPAN>files (the YAML's<SPAN>&nbsp;</SPAN><CODE>sql_file:</CODE><SPAN>&nbsp;</SPAN>points to the path). The YAML captures<SPAN>&nbsp;</SPAN><EM>"which DP is classified where + where it comes from + how it is partitioned + which relationship is realized by which View."</EM><SPAN>&nbsp;</SPAN>The semantics layered on top belong to the Turtle.</P></BLOCKQUOTE><H3 id="organizing-csn-metadata-for-ontology-design" id="toc-hId--931765425">Organizing CSN metadata for ontology design</H3><P class="">The design process starts with<SPAN>&nbsp;</SPAN><STRONG>organizing the CSN (Core Schema Notation) metadata</STRONG>. The CSN specification for each BDC Data Product is available as JSON, but it contains every field and annotation. The full payload is too verbose to feed directly into ontology design.</P><P class="">So for each DP, the following information is distilled into a<SPAN>&nbsp;</SPAN><STRONG>CSN_SUMMARY</STRONG><SPAN>&nbsp;</SPAN>document:</P><UL class=""><LI><STRONG>Main entity</STRONG><SPAN>&nbsp;</SPAN>name and ABAP source name</LI><LI><STRONG>Key columns</STRONG><SPAN>&nbsp;</SPAN>(PK + business key)</LI><LI><STRONG>Relationship links (FK and similar)</STRONG>: which column connects to which entity in which DP</LI><LI><STRONG>Amount / quantity columns</STRONG>: fields that may become SQL aggregation targets</LI><LI><STRONG>KG / SQL classification candidate</STRONG>: master data or transaction</LI></UL><P class="">This task can also be delegated to the LLM. Hand it the CSN JSON files and the blank template together and ask:<SPAN>&nbsp;</SPAN><EM>"Organize each DP's main entity, keys, FKs, and amount columns into a table."</EM><SPAN>&nbsp;</SPAN>A person reads the result and decides<SPAN>&nbsp;</SPAN><EM>"this DP goes into the KG or SQL; this FK is a relationship candidate,"</EM><SPAN>&nbsp;</SPAN>and those decisions are persisted in the next step's<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE>.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":pushpin:">📌</span><SPAN>&nbsp;</SPAN><STRONG>CSN_SUMMARY and<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>have separate roles.</STRONG><SPAN>&nbsp;</SPAN>CSN_SUMMARY is responsible for<SPAN>&nbsp;</SPAN><EM>"a human-readable compression of raw CSN."</EM><SPAN>&nbsp;</SPAN>The<SPAN>&nbsp;</SPAN><STRONG>results of decisions, </STRONG>such as<SPAN>&nbsp;</SPAN><EM>"DP category classification / multi-hop relationships / Class Partition / VT · synonym mapping"</EM><SPAN>&nbsp;</SPAN>— belong in the next step's<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE>. Putting everything in a single document blurs the line between<SPAN>&nbsp;</SPAN><EM>"compressed reference"</EM><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><EM>"decision,"</EM><SPAN>&nbsp;</SPAN>and any decision change requires editing two places.</P></BLOCKQUOTE><H4 id="vibe-coding-authoring-csn_summary" id="toc-hId--1421681937">Vibe Coding: authoring CSN_SUMMARY</H4><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>docs/dp_specs/*_CSN.json</CODE>: per-DP CSN specs (downloaded alongside the BDC Data Product)</LI><LI><CODE>docs/dp_specs/CSN_SUMMARY_TEMPLATE.md</CODE>: blank template (in the workshop materials, or the demo repository)</LI><LI><CODE>virtual_tables.csv</CODE><SPAN>&nbsp;</SPAN>(enriched): the Section 1 VT list + synonym information</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE>: per-DP compressed reference filled in with domain data<UL class=""><LI>Per DP: main entity / key columns / relationship links / amount · quantity columns / tentative KG · SQL classification</LI></UL></LI><LI>Used in: direct input for the Section 2 dp_mapping.yaml draft</LI></UL></BLOCKQUOTE><P class="">This is the most important step in this post. The LLM reads all CSN JSON files and fills in the template. The domain expert's task is to review whether the classification and relationship candidates fit the domain.</P><PRE><CODE>Copy docs/dp_specs/CSN_SUMMARY_TEMPLATE.md to docs/dp_specs/CSN_SUMMARY.md and fill in every &lt;...&gt; placeholder using these input sources: - dp_specs/*_CSN.json: per-DP main entity, key columns, FKs, amount/quantity columns - virtual_tables.csv (with synonym info): match each DP to its available VT For each DP, fill in every section of the TEMPLATE: - Main entity / ABAP source name / column count - Key columns table - Relationship links table (FKs and similar candidate connections to other DPs, raw) - Amount/quantity columns table - Tentative KG / SQL classification + one-line reason ⚠️ Follow the TEMPLATE's format (sections / table columns) exactly. Adjust the domain vocabulary to fit your data. ⚠️ Do NOT include comprehensive decisions such as "DP category / multi-hop relationships / Class Partition" here. Those belong in the next step's dp_mapping.yaml.</CODE></PRE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span>&nbsp;Our demo's<SPAN>&nbsp;</SPAN><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE><SPAN>&nbsp;</SPAN>is the result of this step. From it, a person can see at a glance which DPs are master data, which are transactions, and which relationships are feasible.</P></BLOCKQUOTE><H3 id="moving-csn_summary-into-a-dp_mappingyaml-draft" id="toc-hId--1324792435">Moving CSN_SUMMARY into a<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>draft</H3><P class="">If CSN_SUMMARY was the<SPAN>&nbsp;</SPAN><EM>"per-DP compressed reference"</EM><SPAN>&nbsp;</SPAN>of what each DP looks like,<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>is the<SPAN>&nbsp;</SPAN><STRONG>SSoT in which comprehensive decisions are persisted</STRONG><SPAN>&nbsp;</SPAN>based on it. Decisions intentionally left out of the CSN_SUMMARY, such as<SPAN>&nbsp;</SPAN><EM>"DP category classification / relationship candidates / VT · synonym mapping",</EM>&nbsp;collect here. The same information is then ① classified by DP-level category (master / structural / transaction), ② grouped with each DP's VT / synonym information, and ③ given placeholders for downstream decisions (<CODE>kg_class</CODE>, multi-hop relationship definitions, and so on).</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":pushpin:">📌</span><SPAN>&nbsp;</SPAN><STRONG>CSN_SUMMARY → YAML mapping.</STRONG><SPAN>&nbsp;</SPAN>The following moves from CSN_SUMMARY into the YAML:</P><UL class=""><LI><STRONG>Per-DP relationship links (FKs and similar)</STRONG><SPAN>&nbsp;</SPAN>→ the basis for tentative<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>placement (after Section 2 JOIN verification, some shift to<SPAN>&nbsp;</SPAN><CODE>structural_data</CODE>)</LI><LI><STRONG>Per-DP tentative KG / SQL classification</STRONG><SPAN>&nbsp;</SPAN>→ tentative category values in the YAML (in Section 3, some move to<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE>)</LI><LI><STRONG>VT / synonym mapping from the Section 1 enriched CSV</STRONG><SPAN>&nbsp;</SPAN>→<SPAN>&nbsp;</SPAN><CODE>virtual_table_schema</CODE>,<SPAN>&nbsp;</SPAN><CODE>virtual_table_name</CODE>,<SPAN>&nbsp;</SPAN><CODE>synonym_name</CODE><SPAN>&nbsp;</SPAN>in the YAML</LI><LI><STRONG>Amount / quantity columns</STRONG><SPAN>&nbsp;</SPAN>→ (not persisted in the YAML; the Part 4 SQL tool references CSN_SUMMARY directly)</LI></UL></BLOCKQUOTE><H4 id="vibe-coding-drafting-dp_mappingyaml" id="toc-hId--1814708947">Vibe Coding: drafting dp_mapping.yaml</H4><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE>: the per-DP compressed reference from Section 2</LI><LI><CODE>virtual_tables.csv</CODE><SPAN>&nbsp;</SPAN>(enriched): the Section 1 VT list + synonym information</LI><LI><CODE>ingestion/dp_mapping.template.yaml</CODE>: blank template (in the workshop materials, or the demo repository)</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>draft: all DPs tentatively placed in the<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>category, with each DP's<SPAN>&nbsp;</SPAN><CODE>virtual_table_schema</CODE>, entity list, and<SPAN>&nbsp;</SPAN><CODE>synonym_name</CODE><SPAN>&nbsp;</SPAN>filled in</LI><LI>Used in: Section 2 JOIN verification, Section 3 KG vs SQL classification, Section 4 Class Partition, Section 5 final Ontology</LI></UL></BLOCKQUOTE><P class="">The LLM reads the CSN_SUMMARY and the enriched CSV and fills in the YAML skeleton. A person confirms only that the tentative category placement is correct and that the synonym names are persisted accurately.</P><PRE><CODE>Follow the structure of ingestion/dp_mapping.template.yaml exactly and produce an ingestion/dp_mapping.yaml draft. Input sources: - docs/dp_specs/CSN_SUMMARY.md : per-DP main entity, keys, relationship links, amounts/quantities, tentative KG/SQL classification - virtual_tables.csv (enriched) : VT list + synonym_schema, synonym_name What to fill in: - Place every DP tentatively in master_data (I will reclassify in later steps) - Leave structural_data and transaction_data with their keys present and empty lists - Leave relationships with its key and an empty list (I will fill in relationship definitions during ontology design) - For each DP, fill in name / description / virtual_table_schema / entities[] - entities[*]: entity, synonym_schema, synonym_name, virtual_table_name - Fill in kg_class from the CSN_SUMMARY tentative classification + domain common sense (for example, WorkCenter DP → bdc:WorkCenter) - Fill in metadata such as ord_id and priority from CSN_SUMMARY or the BDC catalog ⚠️ Follow the TEMPLATE's key names, indentation, and category order exactly. ⚠️ Do NOT add a partition (Class Partition) block in this draft — I will add it to the relevant DPs during the Class Partition decision step. ⚠️ Do NOT add items inside relationships in this draft — I will fill them in during ontology design. Leave the key present and the list empty. ⚠️ Do NOT add datatype_properties for any DP in this draft — I will organize them during ontology design once the class hierarchy is finalized. Finally, do not just author the YAML — verify it directly and report the result: - Confirm that yaml.safe_load() parses successfully - Confirm that every DP is placed in exactly one of the three categories (compare against the DP list in CSN_SUMMARY; nothing missing) - Confirm that each DP's synonym_name matches the value in virtual_tables.csv (enriched) - Confirm that entities[*].virtual_table_name is not empty If verification is not clean, briefly explain what went wrong, fix the YAML, and verify again. This is the input for every subsequent step, so a defect here surfaces only much later — validate it thoroughly at this stage.</CODE></PRE><H4 id="dp_mappingyaml-evolves-step-by-step" id="toc-hId--2011222452"><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>evolves step by step</H4><P class="">The draft above is not completed in one pass.<SPAN>&nbsp;</SPAN><STRONG>Immediately after setup, every DP is placed in<SPAN>&nbsp;</SPAN><CODE>master_data</CODE></STRONG><SPAN>&nbsp;</SPAN>(tentative classification), and the decisions made in the later half of Section 2 through Section 5 refine the YAML classification.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":warning:">⚠️</span>The YAML classifies at the<SPAN>&nbsp;</SPAN><STRONG>DP level</STRONG>.<SPAN>&nbsp;</SPAN><EM>"This DP leans toward a relationship-bridging role"</EM><SPAN>&nbsp;</SPAN>is the coarse classification the YAML captures. The finer entity-level decisions, such as<SPAN>&nbsp;</SPAN><EM>"this entity is a KG node, this entity is a bridge, this entity is ignored"</EM><SPAN>&nbsp;</SPAN>— are addressed in detail in Section 4 (Class Partition) and Part 3 (KGR Views).</P></BLOCKQUOTE><TABLE><TBODY><TR><TD><STRONG>Point in time</STRONG></TD><TD><STRONG>YAML change</STRONG></TD><TD><STRONG>Result of which decision</STRONG></TD></TR><TR><TD>Immediately after draft</TD><TD>All 9 DPs in<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>(tentative),<SPAN>&nbsp;</SPAN><CODE>relationships</CODE><SPAN>&nbsp;</SPAN>empty list</TD><TD>"Everything brought in for now"</TD></TR><TR><TD>After Section 2 JOIN verification</TD><TD>ProductionRouting, PurchasingSourceList →<SPAN>&nbsp;</SPAN><CODE>structural_data</CODE><SPAN>&nbsp;</SPAN>(optional category)</TD><TD>Discovered as bridge / junction roles</TD></TR><TR><TD>Section 3 KG vs SQL</TD><TD>SalesOrder, PurchaseOrder, BillingDocument, ProdOrderConfirmation →<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE></TD><TD>Decision: "not in the KG, SQL only"</TD></TR><TR><TD>Section 4 Class Partition</TD><TD><CODE>partition</CODE><SPAN>&nbsp;</SPAN>block added to the Product DP (FERT/HALB/ROH → FinishedGood/SemiFinished/RawMaterial)</TD><TD>"One DP, three classes"</TD></TR><TR><TD>Section 5 final Ontology</TD><TD><CODE>relationships</CODE><SPAN>&nbsp;</SPAN>filled with ObjectProperty definitions + KGR View mappings</TD><TD>Semantic relationships between classes confirmed</TD></TR><TR><TD>Section 6 Turtle authoring</TD><TD>YAML stops changing. ontology.ttl is generated from YAML + separate label/comment</TD><TD>Ready for loading</TD></TR></TBODY></TABLE><P class="">Keep this table in mind as you read Sections 2 (latter half) through 5, and<SPAN>&nbsp;</SPAN><EM>"where in the YAML each decision is persisted"</EM><SPAN>&nbsp;</SPAN>becomes apparent. At the end of each subsection, we briefly summarize the<SPAN>&nbsp;</SPAN><EM>"YAML changes resulting from this step."</EM></P><H3 id="csn-relationship-links-fks-and-similar-are-a-starting-point-bind-business-relationships-on-top-and-verify-with-actual-joins" id="toc-hId--1914332950">CSN relationship links (FKs and similar) are a starting point; bind business relationships on top and verify with actual JOINs</H3><P class="">CSN metadata indicates<SPAN>&nbsp;</SPAN><EM>"this column is the relationship link (FK) of that table"</EM><SPAN>&nbsp;.</SPAN>&nbsp;However, that is only a starting point. The ontology does not end there. Domain analysis adds<SPAN>&nbsp;</SPAN><STRONG>business relationships not present in CSN</STRONG><SPAN>&nbsp;</SPAN>(for example,<SPAN>&nbsp;</SPAN><EM>"which Products a WorkCenter produces,"</EM><SPAN>&nbsp;</SPAN>&nbsp;a relationship the domain requires but that is not directly expressed as a standard relationship link), and<SPAN>&nbsp;</SPAN><STRONG>relationships that span different Data Products</STRONG><SPAN>&nbsp;</SPAN>(for example,<SPAN>&nbsp;</SPAN><CODE>locatedAt</CODE><SPAN>&nbsp;</SPAN>between<SPAN>&nbsp;</SPAN><CODE>WorkCenter</CODE><SPAN>&nbsp;</SPAN>in the WorkCenter DP and<SPAN>&nbsp;</SPAN><CODE>Plant</CODE><SPAN>&nbsp;</SPAN>in the Plant DP) must be bound by a person. CSN reports only relationships within a single DP; cross-DP relationships are filled in with domain knowledge. Whether a relationship comes from CSN, spans DPs, or is a business relationship,<SPAN>&nbsp;</SPAN><STRONG>whether the JOIN actually works on real data</STRONG><SPAN>&nbsp;</SPAN>must be verified separately.</P><P class="">This post follows<SPAN>&nbsp;</SPAN><STRONG>R2RML principles</STRONG><SPAN>&nbsp;</SPAN>when defining and verifying relationships. Regardless of whether the relationship is declared in CSN, as long as the values exist and a domain expert designates it a relationship, it becomes one. For the detailed R2RML standard, consult<SPAN>&nbsp;</SPAN><STRONG>Part 2-1 Section 3</STRONG>.</P><P class="">It is common for metadata to declare a relationship while the data is missing, or for data types to mismatch subtly, or for one side's column to have so many NULLs that the relationship is effectively broken. Missing this during ontology authoring leads to<SPAN>&nbsp;</SPAN><EM>"the relationship was declared, yet SPARQL returns zero rows"</EM><SPAN>&nbsp;</SPAN>after loading.</P><P class=""><STRONG>The verification approach</STRONG>: execute the actual JOINs for both CSN's FK candidates and the business relationship candidates we added, and check the result row counts. If the result is 0 or far fewer than expected, that relationship is dropped from the ontology, or the data source must be revisited.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":light_bulb:">💡</span>Our demo's<SPAN>&nbsp;</SPAN><CODE>ingestion/02_verify_joins.py</CODE><SPAN>&nbsp;</SPAN>performs this work. It collects the JOIN candidates (as SQL JOIN statements) within the script, executes them one by one, and prints the matching row counts. This script revealed that<SPAN>&nbsp;</SPAN><STRONG>every row of<SPAN>&nbsp;</SPAN><CODE>ProductionRoutingHeader.WorkCenterInternalID</CODE><SPAN>&nbsp;</SPAN>in our demo data was<SPAN>&nbsp;</SPAN><CODE>'00000000'</CODE>, breaking the routing connection</STRONG><SPAN>&nbsp;</SPAN>(the full case is in Part 5 Challenge 2), and triggered a redesign of the multi-hop path in the ontology.</P></BLOCKQUOTE><P class="">This verification is repeated every time a new relationship is added to the ontology. The verification work becomes the foundation for the data pipeline in the next post (Part 3). Only verified relationships are persisted as ObjectProperties in the ontology, and the resulting KG actually works.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":pushpin:">📌</span><SPAN>&nbsp;</SPAN><STRONG>This verification report is the basis for the decisions in Sections 3–5.</STRONG><SPAN>&nbsp;</SPAN>Section 3's KG vs SQL classification, Section 4's Class Partition, and Section 5's final ontology structure, all are decisions made after confirming<SPAN>&nbsp;</SPAN><EM>"which relationships are alive in the real data."</EM><SPAN>&nbsp;</SPAN>Without verification, guessing at the ontology leads to SPARQL returning zero rows after loading.</P></BLOCKQUOTE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":ledger:">📒</span><SPAN>&nbsp;</SPAN><STRONG>YAML changes from this step + one domain decision.</STRONG></P><P class="">JOIN verification revealed two DPs that play<SPAN>&nbsp;</SPAN><EM>"relationship / edge roles."</EM><SPAN>&nbsp;</SPAN><CODE>ProductionRouting</CODE><SPAN>&nbsp;</SPAN>(a multi-hop intermediate between WorkCenter and Product) and<SPAN>&nbsp;</SPAN><CODE>PurchasingSourceList</CODE><SPAN>&nbsp;</SPAN>(a Supplier ↔ Product bridge). Both move from<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>to<SPAN>&nbsp;</SPAN><CODE>structural_data</CODE>. One more decision is needed:<SPAN>&nbsp;</SPAN><STRONG>how to handle multi-hop bridges in the KG.</STRONG></P><P class=""><STRONG>Two paths when encountering a multi-hop bridge:</STRONG></P><TABLE><TBODY><TR><TD width="120.57px" height="30px"><STRONG>Option&nbsp;</STRONG></TD><TD width="276.367px" height="30px"><STRONG>How the bridge is handled</STRONG></TD><TD width="223.914px" height="30px"><STRONG>Pros</STRONG></TD><TD width="205.148px" height="30px"><STRONG>Cons</STRONG></TD></TR><TR><TD width="120.57px" height="112px"><STRONG>(a) Direct connect (flatten)</STRONG></TD><TD width="276.367px" height="112px">The bridge entity is not persisted in the KG; only the endpoints (A → C) are connected as a 1-hop relationship</TD><TD width="223.914px" height="112px">The KG stays clean, multi-hop does not pollute visualization, SPARQL is simple</TD><TD width="205.148px" height="112px">Information carried by the bridge itself is invisible in the KG</TD></TR><TR><TD width="120.57px" height="85px"><STRONG>(b) Persist the bridge as a node</STRONG></TD><TD width="276.367px" height="85px">The bridge entity is loaded into the KG as a separate class (A → B → C)</TD><TD width="223.914px" height="85px">The bridge's attributes (plant, period, priority) become explorable via SPARQL</TD><TD width="205.148px" height="85px">Node count increases, intermediate nodes appear in visualizations</TD></TR></TBODY></TABLE><P class=""><STRONG>We chose different answers for the two DPs.</STRONG></P><UL class=""><LI><CODE>ProductionRouting</CODE><SPAN>&nbsp;</SPAN>→<SPAN>&nbsp;</SPAN><STRONG>(a) flatten.</STRONG><SPAN>&nbsp;</SPAN>The routing header / operations themselves carry no attributes worth querying via SPARQL beyond the fact that<SPAN>&nbsp;</SPAN><EM>"this WorkCenter produces this Product."</EM><SPAN>&nbsp;</SPAN>A SQL View that compresses 4-hop into 1-hop (covered in Part 3) is created, and no nodes are loaded into the KG.</LI><LI><CODE>PurchasingSourceList</CODE><SPAN>&nbsp;</SPAN>→<SPAN>&nbsp;</SPAN><STRONG>(b) persist as a node.</STRONG><SPAN>&nbsp;</SPAN>PSL carries information of its own, such as<SPAN>&nbsp;</SPAN><CODE>plant</CODE>,<SPAN>&nbsp;</SPAN><CODE>validity period</CODE>, and<SPAN>&nbsp;</SPAN><CODE>priority</CODE>&nbsp;beyond the basic<SPAN>&nbsp;</SPAN><EM>"this supplier provides this material."</EM><SPAN>&nbsp;</SPAN>Answering questions such as<SPAN>&nbsp;</SPAN><EM>"show only priority-1 sources"</EM><SPAN>&nbsp;</SPAN>requires PSL to be a node.</LI></UL><P class=""><STRONG>Both DPs go into the<SPAN>&nbsp;</SPAN><CODE>structural_data</CODE><SPAN>&nbsp;</SPAN>category</STRONG><SPAN>&nbsp;</SPAN>because both serve<SPAN>&nbsp;</SPAN><EM>"a supporting role connecting the core nodes in master_data."</EM><SPAN>&nbsp;</SPAN>Within that category, whether a node is persisted in the KG (b) depends on whether the domain wants to ask further questions of that bridge. The fate of each is depicted concretely in Section 5 (final Ontology).</P><P class=""><CODE>structural_data</CODE><SPAN>&nbsp;</SPAN>is an optional category that may be left empty depending on the scenario. If the core nodes in master_data alone connect every relationship, there is no need for it. We separated it as its own category because multi-hop is in play and a junction with its own attributes (PSL) is present.</P></BLOCKQUOTE><H4 id="vibe-coding-the-join-verification-script" id="toc-hId-2058901525">Vibe Coding: the JOIN verification script</H4><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>ingestion/02_verify_joins.py</CODE><SPAN>&nbsp;</SPAN>(workshop materials / demo repository): the infrastructure code + JOINS / SAMPLE_QUERIES structure as a reference. Hand the LLM the file and say<SPAN>&nbsp;</SPAN><EM>"following the structure and output format, fill in the JOINS for our domain only."</EM></LI><LI><CODE>ingestion/dp_mapping.yaml</CODE>: synonym mapping (BDC_DP.* short names)</LI><LI><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE>: relationship candidates + KG / SQL classification</LI><LI>Competency Questions (CQs): the business questions the KG must answer</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI><CODE>ingestion/02_verify_joins.py</CODE>: a verification script populated with your domain's JOIN candidates</LI><LI>Verification report: per-candidate row count and match rate (execution stdout / log)</LI><LI>Used in:<SPAN>&nbsp;</SPAN><STRONG>the basis for Section 3 KG vs SQL classification, Section 4 Class Partition, and Section 5 final ontology decisions</STRONG></LI></UL></BLOCKQUOTE><P class="">The verification script mixes work requiring domain knowledge (<EM>"is this relationship business-meaningful"</EM>) with mechanical repetition (<EM>"count the rows of this JOIN"</EM>). Delegating the latter to the LLM lets a person focus exclusively on the former. The prompt should contain four elements:</P><OL class=""><LI><STRONG>Specify the metadata source.</STRONG><SPAN>&nbsp;</SPAN>Tell the LLM where to find candidates. Point it at the<SPAN>&nbsp;</SPAN><CODE>SYS.VIRTUAL_TABLES</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>PUBLIC.VIRTUAL_COLUMNS</CODE><SPAN>&nbsp;</SPAN>metadata views, and hand it<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>(the VT → synonym mapping) so it writes SQL with short names (<CODE>BDC_DP.WORK_CENTER</CODE>). Without the YAML, the LLM guesses long VT names and fails.</LI><LI><STRONG>Provide the Competency Questions (CQs).</STRONG><SPAN>&nbsp;</SPAN>Hand over<SPAN>&nbsp;</SPAN><EM>"the questions the KG must answer"</EM><SPAN>&nbsp;</SPAN>so the LLM can narrow the candidates by<SPAN>&nbsp;</SPAN><EM>"which relationship resolves this question"</EM><SPAN>&nbsp;</SPAN>(for example,<SPAN>&nbsp;</SPAN><EM>"who is affected by a WorkCenter shutdown?"</EM><SPAN>&nbsp;</SPAN>→ the WorkCenter ↔ Product ↔ SalesOrder ↔ Customer chain becomes the verification target).</LI><LI><STRONG>Concrete verification criteria.</STRONG><SPAN>&nbsp;</SPAN>Replace vague language such as<SPAN>&nbsp;</SPAN><EM>"works well"</EM><SPAN>&nbsp;</SPAN>with numeric thresholds: row counts, NULL ratios, minimum matching thresholds.</LI><LI><STRONG>Human-reviewable output format.</STRONG><SPAN>&nbsp;</SPAN>Tabular format + sample rows.<SPAN>&nbsp;</SPAN><EM>"What you can see with your eyes"</EM><SPAN>&nbsp;</SPAN>becomes the input for the domain decision.</LI></OL><PRE><CODE>Good prompt: "Following the structure and output format of ingestion/02_verify_joins.py exactly, rewrite ingestion/02_verify_joins.py for our domain. Replace only the JOINS and SAMPLE_QUERIES lists; do NOT modify the infrastructure functions (run_join_checks, run_sample_queries, main). [Input materials / attach together] - 02_verify_joins.py (existing) : structure and output format reference - dp_mapping.yaml : VT → synonym mapping (BDC_DP.* short names) - (if needed) CSN_SUMMARY.md : relationship candidates (per-DP reference) Competency Questions the KG must answer: - CQ1: When a specific WorkCenter goes down, which customers' revenue is affected? - CQ2: Which Product is produced at which Plant? - CQ3: Which Supplier supplies which Product? ... (as many as needed) First organize which relationship chains are needed to answer these CQs, and on top of those chains, populate the JOINS list with three kinds of candidates: (a) Direct relationship links within one DP : identify candidates by column name patterns (for example, '*ID', '*Code') (b) Relationships across DPs : the same column name or the same domain values appearing in multiple DPs (c) Multi-hop business relationships : chains added through domain analysis (I will provide the JOIN SQL separately) Write SQL with the short synonym names in dp_mapping.yaml (BDC_DP.WORK_CENTER and similar). Place each entry in the same tuple form as the existing file (label, key columns description, SQL). Additionally, populate SAMPLE_QUERIES with 3–5 spot-check queries that a person can verify by eye (for example, 'top 10 WorkCenters', 'sample 5 rows joined via a specific relationship'). Finally, do not just author the script — execute it and show me the results. - If every JOIN matches with ✅, you are done. - If any item returns ⚠️ 0 rows or ⛔ an error, briefly hypothesize the cause (wrong column name? missing data? a different path needed?) and propose a fix. Once I approve, modify JOINS and re-run. Repeat until the result is clean. [After verification passes, persist candidates to the YAML] Once verification is clean, persist the surviving JOINs into the relationships section of dp_mapping.yaml as 'free-form candidates.' Do not yet populate the formal ObjectProperty form (name/domain/range/source) — I will organize that during ontology design. Example (one item each, free-form description only): relationships: - description: "WorkCenter → ProductionRoutingOpSubord (key: WorkCenterInternalID+TypeCode, matched: 8,432 rows)" - description: "Product → ProductPlant via Material (matched: 3,197 rows)" - ... ⚠️ Do NOT touch master_data / structural_data / transaction_data / partition / datatype_properties. ⚠️ After modifying the YAML, confirm yaml.safe_load() parses successfully."</CODE></PRE><P class="">→ Once the LLM organizes<SPAN>&nbsp;</SPAN><EM>"which relationship chains are needed to answer the CQs"</EM><SPAN>&nbsp;</SPAN>and writes the verification script, a person reads the results and decides<SPAN>&nbsp;</SPAN><STRONG>"keep this relationship, drop that one."</STRONG><SPAN>&nbsp;</SPAN>Our demo's<SPAN>&nbsp;</SPAN><CODE>02_verify_joins.py</CODE><SPAN>&nbsp;</SPAN>was built exactly this way, so handing that file as a reference lets the LLM fill in JOINs for your own domain in the same structure and format.</P><PRE><CODE>Bad prompt: "Verify that the FKs work properly." → "Properly" has no criterion. The LLM looks at the metadata only and says "FK is declared, so OK." → Without the CQs the KG must answer, no way to distinguish important relationships. → Without the synonym list, the LLM guesses long VT names or writes generic SQL. → Cross-DP and multi-hop relationships are out of sight entirely.</CODE></PRE><P class=""><STRONG>Roles:</STRONG></P><UL class=""><LI><STRONG>LLM</STRONG>: drafting<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>+ step-by-step updates (moving DPs between<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>structural_data</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE><SPAN>&nbsp;</SPAN>based on verification results, filling in<SPAN>&nbsp;</SPAN><CODE>kg_class</CODE>), writing catalog query SQL, deriving relationship chains from the CQs, writing JOIN execution scripts, organizing result tables, flagging clearly broken relationships (zero matches)</LI><LI><STRONG>Domain expert</STRONG>: ① defining CQs (the list of business questions the KG must answer), ② reviewing whether the LLM's proposed YAML classification and naming match the domain vocabulary, ③ deciding which relationship candidates to add to the verification target (especially cross-DP and multi-hop), ④ judging whether to keep or drop relationships with low match rates, ⑤ analyzing whether broken relationships are due to missing data or whether they can be routed differently</LI></UL><P class="">Section 2 deliverables:</P><TABLE><TBODY><TR><TD><STRONG>Step Deliverable Role</STRONG></TD><TD><STRONG>Step Deliverable</STRONG></TD><TD><STRONG>Role</STRONG></TD></TR><TR><TD>CSN_SUMMARY authoring</TD><TD><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE></TD><TD>Per-DP main entity / keys / relationship links (FKs and similar) / amounts · quantities (compressed per DP)</TD></TR><TR><TD><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>draft</TD><TD><CODE>ingestion/dp_mapping.yaml</CODE></TD><TD>SSoT draft with category classification, VT / synonym mapping, and<SPAN>&nbsp;</SPAN><CODE>kg_class</CODE><SPAN>&nbsp;</SPAN>placeholders filled in</TD></TR><TR><TD>JOIN verification script</TD><TD><CODE>ingestion/02_verify_joins.py</CODE><SPAN>&nbsp;</SPAN>+ verification report</TD><TD>Match rate and NULL ratio per relationship candidate,<SPAN>&nbsp;</SPAN><STRONG>the basis for decisions in Sections 3–5</STRONG></TD></TR></TBODY></TABLE><H1 id="3-deciding-between-kg-and-sql" id="toc-hId--1552370255">3. Deciding Between KG and SQL</H1><P class="">After working through Section 2, the categorization in<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>is refined once. Core node DPs are in<SPAN>&nbsp;</SPAN><CODE>master_data</CODE>, DPs that emerged as bridges are in<SPAN>&nbsp;</SPAN><CODE>structural_data</CODE>, and the<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE><SPAN>&nbsp;</SPAN>slot is still empty.</P><P class="">This section makes two decisions:</P><OL class=""><LI><STRONG>Section 3.1</STRONG>: transactional data (numeric aggregation targets) go to SQL, moving to<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE></LI><LI><STRONG>Section 3.2</STRONG>: master data orphaned by those moves (entities that became "islands" in the KG) also go to SQL</LI></OL><P class="">After these two decisions, only<SPAN>&nbsp;</SPAN><EM>"master data whose structural relationships must be traversed"</EM><SPAN>&nbsp;</SPAN>remains in the KG.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>Entities whose structural relationships must be traversed → KG. Data that must be aggregated numerically → SQL. Entities that are not connected to other nodes within the KG → SQL.</STRONG></P></BLOCKQUOTE><P class="">Our final classification (instance counts below are SAP demo data examples):</P><PRE><CODE>KG (master data, structural relationships): ├── WorkCenter , work center (129 examples) ├── Product , finished / semi-finished / raw material (2,593 examples) ├── Supplier , supplier (153 examples) ├── Plant , plant (12 examples) └── PurchasingSourceList ← Supplier ↔ Product bridge node SQL (transactional, numeric): ├── SalesOrder / SalesOrderItem ├── PurchaseOrder / PurchaseOrderItem ├── BillingDocument / BillingDocumentItem └── ProdOrderConfirmation SQL (master orphaned as an island, lookup only): └── Customer → name / country lookup via SQL CUSTOMER synonym</CODE></PRE><P class="">The W3C standards do not make this decision. The standard only says<SPAN>&nbsp;</SPAN><EM>"create reference triples when FKs exist."</EM><SPAN>&nbsp;</SPAN><EM>"Should this entity live in the KG?"</EM><SPAN>&nbsp;</SPAN>is a domain judgment.</P><H2 id="31-transactional-data-goes-to-sql" id="toc-hId--2042286767">3.1 Transactional data goes to SQL</H2><P class="">The first decision is clear.<SPAN>&nbsp;</SPAN><EM>"Data whose core is dates, amounts, and counts"</EM><SPAN>&nbsp;</SPAN>does not belong in the KG.</P><P class="">Reasons not to load it into the KG:</P><UL class=""><LI><STRONG>Instance explosion</STRONG>: transactions are voluminous and change frequently. SalesOrder alone runs into tens or hundreds of thousands of rows. Graph nodes increase in proportion, weighing down both visualization and SPARQL query cost.</LI><LI><STRONG>Numbers are natural in SQL</STRONG>: aggregations such as<SPAN>&nbsp;</SPAN><EM>"revenue per customer"</EM><SPAN>&nbsp;</SPAN>are much simpler and faster with SQL<SPAN>&nbsp;</SPAN><CODE>GROUP BY</CODE><SPAN>&nbsp;</SPAN>than with SPARQL.</LI><LI><STRONG>Not a KG question</STRONG>: the KG exists to answer<SPAN>&nbsp;</SPAN><EM>"what does this WorkCenter produce?"</EM><SPAN>&nbsp;</SPAN>style structural traversals, not<SPAN>&nbsp;</SPAN><EM>"what is the total?"</EM></LI></UL><P class="">In our scenario, four DPs match these criteria:</P><TABLE><TBODY><TR><TD><STRONG>DP</STRONG></TD><TD><STRONG>Why it is transactional</STRONG></TD></TR><TR><TD><CODE>SalesOrder</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>SalesOrderItem</CODE></TD><TD>Sale date · amount · per-customer orders</TD></TR><TR><TD><CODE>PurchaseOrder</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>PurchaseOrderItem</CODE></TD><TD>Purchase date · amount · per-supplier orders</TD></TR><TR><TD><CODE>BillingDocument</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>BillingDocumentItem</CODE></TD><TD>Billing date · amount</TD></TR><TR><TD><CODE>ProductionOrderConfirmation</CODE></TD><TD>Production completion date · quantity</TD></TR></TBODY></TABLE><P class="">These four DPs move from<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>to<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE>. The actual YAML update is handled together with the Section 3.2 island candidate check via a single Vibe Coding session at the end of Section 3.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":ledger:">📒</span><SPAN>&nbsp;</SPAN><STRONG>YAML changes from this step:</STRONG><SPAN>&nbsp;</SPAN>the four DPs above move from<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>to<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE>, and each DP's<SPAN>&nbsp;</SPAN><CODE>aggregation_columns</CODE><SPAN>&nbsp;</SPAN>is populated. Their entities are not persisted as KG nodes; they are used only as SQL aggregation targets. They remain in the YAML because the synonyms themselves are still needed for SQL queries (classification differs, but the synonyms persist).</P></BLOCKQUOTE><H2 id="32-then-master-data-that-became-an-island-also-goes-to-sql" id="toc-hId-2056167024">3.2 Then master data that became an island also goes to SQL</H2><P class="">After removing the transactional DPs in Section 3.1, looking again at the KG reveals<SPAN>&nbsp;</SPAN><STRONG>an unexpected side effect</STRONG>. A piece of master data has become<SPAN>&nbsp;</SPAN><EM>"an island."</EM></P><P class="">In our scenario,<SPAN>&nbsp;</SPAN><CODE>Customer</CODE><SPAN>&nbsp;</SPAN>did. Customer is master data, and the demo data has roughly 215 instances. Initially, we naturally intended to place it in the KG.<SPAN>&nbsp;</SPAN><EM>"Customers who buy Products"</EM><SPAN>&nbsp;</SPAN>is a core relationship in the scenario.</P><P class="">But there was a problem.<SPAN>&nbsp;</SPAN><STRONG>The bridge connecting Product and Customer is SalesOrder, and we removed SalesOrder from the KG in Section 3.1.</STRONG><SPAN>&nbsp;</SPAN>The KG then looks like this:</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="WorkCenter to FinishedGood-2026-07-03-194431.svg" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429081i2451F49067AA0129/image-size/large?v=v2&amp;px=999" role="button" title="WorkCenter to FinishedGood-2026-07-03-194431.svg" alt="WorkCenter to FinishedGood-2026-07-03-194431.svg" /></span></P><P class="">With the bridge (SalesOrder) absent from the KG, the KG path from Product to Customer is severed. Even if Customer is persisted in the KG, it is<SPAN>&nbsp;</SPAN><STRONG>an isolated island</STRONG><SPAN>&nbsp;</SPAN>unconnected to other nodes. The Agent cannot navigate outward from a Customer instance via SPARQL, and other KG nodes cannot reach Customer. Roughly 200 instances would remain as an isolated collection of nodes.</P><P class="">So we decided to<SPAN>&nbsp;</SPAN><EM>"handle the Product → Customer connection in SQL"</EM><SPAN>&nbsp;</SPAN>(since SalesOrder is already in SQL, one SQL JOIN on<SPAN>&nbsp;</SPAN><CODE>SALES_ORDER_ITEM.Material</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>SoldToParty</CODE><SPAN>&nbsp;</SPAN>suffices), and moved Customer itself to a SQL lookup. Looking up name and country by ID is all that is required.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG>Even for master data, if it is not connected to other nodes within the KG, there is little reason to place it there.</STRONG><SPAN>&nbsp;</SPAN>For simple lookup, SQL is simpler.</P></BLOCKQUOTE><H3 id="island-candidate-signals" id="toc-hId-1566250512">"Island candidate" signals</H3><P class="">To make the same judgment in your own domain, the master data that will become<SPAN>&nbsp;</SPAN><EM>"an island"</EM><SPAN>&nbsp;</SPAN>becomes visible only after working through Section 3.1. Two signals:</P><OL class=""><LI><STRONG>Is the bridge transactional, so it moved to SQL?</STRONG><SPAN>&nbsp;</SPAN>If the only path connecting this master to other masters lies inside a transactional DP that moved to SQL in Section 3.1, placing it in the KG produces a disconnected node. (The DPs the LLM reports as<SPAN>&nbsp;</SPAN><EM>"island candidates"</EM><SPAN>&nbsp;</SPAN>at the end of the Section 3.1 Vibe Coding fall here.)</LI><LI><STRONG>Is lookup sufficient?</STRONG><SPAN>&nbsp;</SPAN>If there is no<SPAN>&nbsp;</SPAN><EM>"navigating outward"</EM><SPAN>&nbsp;</SPAN>from this master to other entities via SPARQL (only looking up names and attributes by ID), a SQL lookup is simpler.</LI></OL><P class="">When both apply, moving to SQL is the natural choice.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":warning:">⚠️</span><SPAN>&nbsp;</SPAN><STRONG>Limits of island detection: at design time, candidates only; final confirmation comes after running the Agent.</STRONG><SPAN>&nbsp;</SPAN>The Section 3.1 Vibe Coding catches the structural signal<SPAN>&nbsp;</SPAN><EM>"the bridge moved to SQL"</EM><SPAN>&nbsp;</SPAN>automatically, but the usage signal<SPAN>&nbsp;</SPAN><EM>"SPARQL never actually starts here"</EM><SPAN>&nbsp;</SPAN>surfaces only after running the Agent. We too initially placed Customer in the KG, then after running the Z_ASM3 scenario and confirming that<SPAN>&nbsp;</SPAN><EM>"it was never used as a starting point,"</EM><SPAN>&nbsp;</SPAN>moved it to SQL. The detailed case is covered in Part 5 Challenge.</P></BLOCKQUOTE><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":ledger:">📒</span><SPAN>&nbsp;</SPAN><STRONG>YAML changes from this step:</STRONG><SPAN>&nbsp;</SPAN>masters confirmed as islands among the<SPAN>&nbsp;</SPAN><EM>"island candidates"</EM><SPAN>&nbsp;</SPAN>reported in Section 3.1 (in our case, Customer) drop out of the KG node set. They remain in the<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>category in the YAML, but they are not loaded as KG classes when authoring the Section 5 final ontology Turtle.</P></BLOCKQUOTE><P>&nbsp;</P><H3 id="vibe-coding-handling-section-3-in-one-shot-transaction-classification--island-candidate-reporting" id="toc-hId-1369737007">Vibe Coding: handling Section 3 in one shot (transaction classification + island candidate reporting)</H3><P class="">The decisions of Section 3.1 and Section 3.2 are handled together in a single LLM prompt. Immediately after classifying transactional DPs, ask the LLM to check whether<SPAN>&nbsp;</SPAN><EM>"any master data depended on a bridge that was just moved to transactional,"</EM><SPAN>&nbsp;</SPAN>and the Section 3.2 review goes much faster.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>(updated through Section 2 —<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>+<SPAN>&nbsp;</SPAN><CODE>structural_data</CODE>)</LI><LI><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE><SPAN>&nbsp;</SPAN>(per-DP amount and quantity columns documented)</LI><LI>CQ list</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>updated: transactional DPs moved to<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE><SPAN>&nbsp;</SPAN>+ each DP's<SPAN>&nbsp;</SPAN><CODE>aggregation_columns</CODE><SPAN>&nbsp;</SPAN>populated</LI><LI><STRONG>Island candidate report</STRONG>: the list of master DPs (if any) disconnected from the KG as a result of the transactional moves</LI></UL></BLOCKQUOTE><PRE><CODE>Examine ingestion/dp_mapping.yaml, docs/dp_specs/CSN_SUMMARY.md, and the CQ list, and pick the "DPs that should be classified as transactional," moving them into the transaction_data category. Transaction criteria: - DPs whose core is dates, amounts, and counts (consult the amount/quantity columns in CSN_SUMMARY) - High volume, frequently changing (SalesOrder, PurchaseOrder, BillingDocument and similar) - Appear in the CQs in "total / average / aggregation" form (numeric aggregation rather than structural traversal) When moving: - For each DP, populate aggregation_columns with the amount/quantity columns from CSN_SUMMARY - Do NOT touch the master_data / structural_data classification (decided in earlier steps) ⚠️ Important: after moving transactional DPs, examine master_data once more. If a master DP's only path to other master nodes ran through a transactional DP that was just moved, that master will become an "island" in the KG. If such a candidate exists, report it separately as a "KG island candidate" (leave the YAML category as master_data for now — I will move it after confirmation). Finally: - Which DPs moved to transaction_data + what aggregation_columns were populated - If there are island candidates, which DP is suspected and why (which transactional DP was the bridge) Summarize briefly and show me. ⚠️ After modifying the YAML, always confirm yaml.safe_load() parses successfully, and if it breaks, briefly explain what went wrong, fix it, and verify again.</CODE></PRE><P class="">A person reviews the LLM's<SPAN>&nbsp;</SPAN><EM>"island candidates"</EM><SPAN>&nbsp;</SPAN>and confirms which ones to remove from the KG node set (the YAML category stays as<SPAN>&nbsp;</SPAN><CODE>master_data</CODE>; the entity simply is not persisted as an ontology class in Section 5).</P><HR /><P>Section 3 deliverables:</P><TABLE><TBODY><TR><TD><STRONG>Step</STRONG></TD><TD><STRONG>Deliverable</STRONG></TD><TD><STRONG>Role</STRONG></TD></TR><TR><TD>Transactional → SQL (Section 3.1)</TD><TD><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>update (4 DPs to<SPAN>&nbsp;</SPAN><CODE>transaction_data</CODE>)</TD><TD>SQL aggregation targets confirmed</TD></TR><TR><TD>Islands → SQL (Section 3.2)</TD><TD><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>update (island candidate entities excluded from KG nodes)</TD><TD>Entities for the KG and entities for SQL-only confirmed</TD></TR></TBODY></TABLE><H1 id="4-splitting-one-data-product-into-multiple-classes-class-partition" id="toc-hId-1760029516">4. Splitting One Data Product into Multiple Classes (Class Partition)</H1><P class="">After the KG vs SQL classification, the next frequently encountered domain decision is<SPAN>&nbsp;</SPAN><STRONG>Class Partition</STRONG>.</P><P class="">SAP BDC has a single Product Data Product. HANA Cloud has a single corresponding Virtual Table (<CODE>BDC_DP.PRODUCT</CODE>). Inside that table are finished goods, semi-finished goods, and raw materials. The distinguishing field is<SPAN>&nbsp;</SPAN><CODE>MaterialType</CODE>.</P><PRE><CODE><SPAN class="">SELECT</SPAN> "MaterialType", <SPAN class="">COUNT</SPAN>(<SPAN class="">*</SPAN>) <SPAN class="">AS</SPAN> cnt <SPAN class="">FROM</SPAN> BDC_DP.PRODUCT <SPAN class="">GROUP</SPAN> <SPAN class="">BY</SPAN> "MaterialType" <SPAN class="">ORDER</SPAN> <SPAN class="">BY</SPAN> cnt <SPAN class="">DESC</SPAN>;</CODE></PRE><P>&nbsp;</P><TABLE width="335px"><TBODY><TR><TD width="113.109px"><STRONG>MaterialType</STRONG></TD><TD width="220.891px"><STRONG>SAP meaning</STRONG></TD></TR><TR><TD width="113.109px"><CODE>HALB</CODE></TD><TD width="220.891px">Semi-finished (Halbfabrikat)</TD></TR><TR><TD width="113.109px"><CODE>FERT</CODE></TD><TD width="220.891px">Finished good (Fertigprodukt)</TD></TR><TR><TD width="113.109px"><CODE>ROH</CODE></TD><TD width="220.891px">Raw material (Rohmaterial)</TD></TR></TBODY></TABLE><P class="">Applied as-is, W3C Direct Mapping would map all three to a single<SPAN>&nbsp;</SPAN><CODE>bdc:Product</CODE>. The data is stored. But we split it into three classes for two reasons.</P><H3 id="reason-1-the-directionality-of-relationships-differs" id="toc-hId-976709997">Reason 1: the directionality of relationships differs</H3><P class="">In a supply chain BOM, the flow is always one-directional.</P><PRE><CODE>WorkCenter ──producesProduct──▶ FinishedGood FinishedGood ──hasComponent──▶ SemiFinished SemiFinished ──hasComponent──▶ RawMaterial</CODE></PRE><P class="">If all three are<SPAN>&nbsp;</SPAN><CODE>bdc:Product</CODE>, a SPARQL query for<SPAN>&nbsp;</SPAN><EM>"what does Z_ASM3 produce?"</EM><SPAN>&nbsp;</SPAN>may return a mix of raw materials and semi-finished goods. Type constraints enforce the correct answer.</P><H3 id="reason-2-an-llm-reasons-from-class-names" id="toc-hId-780196492">Reason 2: an LLM reasons from class names</H3><P class="">When the Agent sees<SPAN>&nbsp;</SPAN><CODE>?p a bdc:FinishedGood</CODE><SPAN>&nbsp;</SPAN>in a SPARQL result, it knows immediately: this is a sellable end product, so look at sales orders and customers. For<SPAN>&nbsp;</SPAN><CODE>bdc:RawMaterial</CODE>, look at procurement, suppliers, and purchase orders.<SPAN>&nbsp;</SPAN><STRONG>The class name conveys business meaning, and the LLM uses that meaning to decide its next action.</STRONG></P><H3 id="the-class-partition-pattern" id="toc-hId-583682987">The Class Partition pattern</H3><P class="">In ontology design, this situation is called<SPAN>&nbsp;</SPAN><STRONG>Class Partition</STRONG><SPAN>&nbsp;</SPAN>or<SPAN>&nbsp;</SPAN><STRONG>Discriminator-based Class Assignment</STRONG>. The pattern assigns different KG classes to rows of a single source (table) based on the value of a specific column (discriminator).</P><P class="">In OWL, this disjointness between classes can be declared explicitly.</P><PRE><CODE># The three classes are mutually disjoint (a finished good cannot also be a raw material) bdc:FinishedGood owl:disjointWith bdc:SemiFinished . bdc:FinishedGood owl:disjointWith bdc:RawMaterial . bdc:SemiFinished owl:disjointWith bdc:RawMaterial . # Or, with owl:disjointUnion (the three classes form a complete partition of Product) bdc:Product owl:disjointUnion (bdc:FinishedGood bdc:SemiFinished bdc:RawMaterial) .</CODE></PRE><H3 id="signals-that-class-partition-should-be-considered" id="toc-hId-555353173">Signals that Class Partition should be considered</H3><P class="">While exploring a Virtual Table, if any of these patterns appears, consider partitioning:</P><UL class=""><LI><STRONG>A type / category column exists</STRONG>:<SPAN>&nbsp;</SPAN><CODE>MaterialType</CODE>,<SPAN>&nbsp;</SPAN><CODE>DocumentType</CODE>,<SPAN>&nbsp;</SPAN><CODE>OrderType</CODE>, and so on</LI><LI><STRONG>Different row sets connect to different entities</STRONG>: some rows link to Supplier, others to Customer</LI><LI><STRONG>Business users say "those are different things"</STRONG>:<SPAN>&nbsp;</SPAN><EM>"finished goods and raw materials are different."</EM></LI></UL><P class="">This pattern appears often in BDC Data Products because S/4HANA table design relies on type columns for flexibility, storing multiple concepts in a single table. ODM surfaces these as meaningful names, but<SPAN>&nbsp;</SPAN><STRONG>whether to split them into KG classes is the designer's decision.</STRONG></P><BLOCKQUOTE dir="auto"><P class=""><STRONG>Design rule: if different row sets participate in different relationships, split a single source table into multiple KG classes.</STRONG></P></BLOCKQUOTE><H3 id="the-class-partition-decision-is-made-by-a-person" id="toc-hId-358839668">The Class Partition decision is made by a person</H3><P class="">Do not delegate<SPAN>&nbsp;</SPAN><EM>"split this table into classes"</EM><SPAN>&nbsp;</SPAN>to the LLM wholesale. The LLM does not know the business meaning and will partition the wrong place, or merge a place that should be partitioned. Which of the three signals above (type column · different relationships per row set · business users distinguishing them) applies to your domain is something only a person knows.</P><H3 id="vibe-coding-class-partition-two-steps" id="toc-hId-162326163">Vibe Coding: Class Partition (two steps)</H3><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE>: DP list under review + synonym (<CODE>BDC_DP.&lt;NAME&gt;</CODE>) and virtual_table_schema / name mapping</LI><LI><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE>: per-DP keys / relationship links / amount · quantity columns (supplementary)</LI><LI>HANA Cloud's<SPAN>&nbsp;</SPAN><CODE>PUBLIC.VIRTUAL_COLUMNS</CODE><SPAN>&nbsp;</SPAN>(column metadata) +<SPAN>&nbsp;</SPAN><CODE>BDC_DP.&lt;SHORT_NAME&gt;</CODE><SPAN>&nbsp;</SPAN>(row queries)</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI>Step 1: partition candidate table (DP · column · distinct values · row counts · expected connections) + first-pass recommendation</LI><LI>Step 2:<SPAN>&nbsp;</SPAN><CODE>partition</CODE><SPAN>&nbsp;</SPAN>block added to the relevant DP in<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>(discriminator + subclass mapping)</LI><LI>In the next step (Section 6), when authoring ontology.ttl, this partition information is converted into Turtle class definitions</LI></UL></BLOCKQUOTE><PRE><CODE>Good prompt / Step 1: find partition candidates "For the DPs in the master_data and structural_data categories of dp_mapping.yaml, organize type/category partition candidates. [Input materials] - ingestion/dp_mapping.yaml : DP list under review + synonym / VT name mapping - docs/dp_specs/CSN_SUMMARY.md : supplementary — per-DP keys / relationship links / amount · quantity columns Each DP entity has two names in the YAML, each with a different purpose: - synonym (for example, BDC_DP.WORK_CENTER) : short name for SQL row/value queries - virtual_table_schema + virtual_table_name : original names for catalog queries ⚠️ HANA Cloud has no DESCRIBE. Column information must be queried from PUBLIC.VIRTUAL_COLUMNS using virtual_table_schema/virtual_table_name (not the synonym). What to organize (automatable): 1. The column list for each entity (query PUBLIC.VIRTUAL_COLUMNS) 2. Candidate columns that look like a type/category (MaterialType, DocumentType, OrderType, and so on) 3. For each candidate column, the distinct values + per-value row counts (SELECT &lt;candidate&gt;, COUNT(*) FROM BDC_DP.&lt;SHORT_NAME&gt; GROUP BY &lt;candidate&gt; — using the synonym) 4. For each distinct value, guess which other DP/column it likely connects to, and run a light feasibility check - The guess is based on column names and domain common sense (for example, MaterialType='FERT' → SalesOrderItem.Material likely) - Confirm only by comparing distinct counts on both columns. Do not run JOINs. If the distinct values on both columns overlap, report "likely connection." Show the result as a table: DP | candidate column | distinct values | per-value row count | expected connected DP·column | expected column's distinct count (+overlap signal). After the table, add a single paragraph with a first-pass recommendation: "Partitioning looks worthwhile for candidates X and Y. Reasons: - X: all N distinct values look meaningful, per-value row counts are balanced, and the connected DPs differ by value (overlap pattern varies by value) - Y: ... Conversely, A and B look weak (for example, a single value covers 99%+, or distinct count is too high suggesting an ID rather than a type column, or every value has only 1–2 rows so domain meaning is weak)." ⚠️ SQL usage limit: stop at distinct + COUNT. Do not run JOINs. - JOINs are not needed to find partition candidates (distinct comparison on both sides is enough) - BDC Virtual Tables are external system calls, so heavy JOINs are costly."</CODE></PRE><P class="">→ Once the LLM produces the table and first-pass recommendation, the next stage is<SPAN>&nbsp;</SPAN><STRONG>manual review by a person</STRONG>.</P><H4 id="manual-review-do-the-type-values-play-different-business-roles" id="toc-hId--327590349">Manual review: do the type values play different business roles?</H4><P class="">The LLM output above (column · distinct values · row counts · expected connections) shows only<SPAN>&nbsp;</SPAN><EM>"the shape of a partition candidate."</EM><SPAN>&nbsp;</SPAN>Judging the real partition value requires one more step.</P><P class=""><STRONG>The core question</STRONG>: does each type value play<SPAN>&nbsp;</SPAN><EM>"a different business role"</EM>? There are two ways to look at this.</P><OL class=""><LI><STRONG>Cross-reference distinct + COUNT results (light, default)</STRONG>: examine the<SPAN>&nbsp;</SPAN><EM>"expected connected DP · column"</EM><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><EM>"expected column's distinct count (+overlap signal)"</EM><SPAN>&nbsp;</SPAN>in the LLM table, and check whether the connections differ per type value. Since this is based only on distinct and COUNT, it is light, and the BDC Virtual Table cost is small.</LI><LI><STRONG>Precise check with JOINs (only when needed)</STRONG>: only when distinct · COUNT comparison cannot reveal the difference. For example,<SPAN>&nbsp;</SPAN><EM>"both columns are 4-character codes with similar distinct counts; I want the actual matching row count"</EM><SPAN>&nbsp;</SPAN>or<SPAN>&nbsp;</SPAN><EM>"one side has trailing spaces, so distinct counts look identical but the JOIN match rate may be low."</EM><SPAN>&nbsp;</SPAN>In these cases, use a verification script such as<SPAN>&nbsp;</SPAN><CODE>02_verify_joins.py</CODE><SPAN>&nbsp;</SPAN>to confirm match rates and NULL ratios directly.</LI></OL><P class="">Example: examining the<SPAN>&nbsp;</SPAN><CODE>MaterialType</CODE><SPAN>&nbsp;</SPAN>candidate of the Product DP</P><P>&nbsp;</P><TABLE><TBODY><TR><TD><STRONG>MaterialType</STRONG></TD><TD><STRONG>How it is referenced in other DPs</STRONG></TD><TD><STRONG>Business role</STRONG></TD></TR><TR><TD><CODE>FERT</CODE><SPAN>&nbsp;</SPAN>(finished good)</TD><TD>Large distinct overlap with<SPAN>&nbsp;</SPAN><CODE>SalesOrderItem.Material</CODE></TD><TD><STRONG>Sold</STRONG></TD></TR><TR><TD><CODE>ROH</CODE><SPAN>&nbsp;</SPAN>(raw material)</TD><TD>Large distinct overlap with<SPAN>&nbsp;</SPAN><CODE>PurchaseOrderItem.Material</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>PurchasingSourceList.Material</CODE></TD><TD><STRONG>Purchased</STRONG></TD></TR><TR><TD><CODE>HALB</CODE><SPAN>&nbsp;</SPAN>(semi-finished)</TD><TD>Appears in<SPAN>&nbsp;</SPAN><CODE>Product.hasComponent</CODE><SPAN>&nbsp;</SPAN>(internal BOM)</TD><TD><STRONG>Internal assembly</STRONG></TD></TR></TBODY></TABLE><P class="">If the three values show<SPAN>&nbsp;</SPAN><STRONG>different connection patterns</STRONG><SPAN>&nbsp;</SPAN>→ different business roles → partition is worthwhile. If the three values show<SPAN>&nbsp;</SPAN><STRONG>the same connection pattern</STRONG><SPAN>&nbsp;</SPAN>→ effectively the same role → partition value is weak, a single class suffices.</P><P class="">This judgment cannot be made by the LLM.<SPAN>&nbsp;</SPAN><STRONG>With the LLM result table in hand</STRONG>, a person decides<SPAN>&nbsp;</SPAN><EM>"are these type values really different roles?"</EM><SPAN>&nbsp;</SPAN>Our scenario's Product followed exactly this analysis and led to the decision<SPAN>&nbsp;</SPAN><EM>"three roles are clearly different → split into three classes."</EM></P><PRE><CODE>Good prompt / Step 2: persist the decision as a YAML partition "Take the Step 1 results and the partition decision I confirmed, and add a partition block to the relevant DP in dp_mapping.yaml. Confirmed partition (example): - DP: BDC_DP.PRODUCT - discriminator column: MaterialType - value → class mapping: HALB → bdc:SemiFinished FERT → bdc:FinishedGood ROH → bdc:RawMaterial Deliverable: partition block added to dp_mapping.yaml - A partition block at the same level as the relevant DP entity (per the template): - name: Product kg_class: bdc:Product entities: - entity: Product synonym_name: PRODUCT ... partition: discriminator: MaterialType parent_class: bdc:Product disjoint: true subclasses: - value: FERT kg_class: bdc:FinishedGood description: "Finished good" - value: HALB kg_class: bdc:SemiFinished description: "Semi-finished" - value: ROH kg_class: bdc:RawMaterial description: "Raw material" - Output form: show the updated DP block as it should appear in the current YAML ⚠️ Do NOT create Turtle class definitions (rdfs:label, rdfs:comment, rdfs:subClassOf, owl:disjointWith) in this step. The ontology.ttl authoring step will take the YAML partition as input and generate them in one pass. ⚠️ In this step, persist only the partition block. Do not touch other regions of master_data (entities, datatype_properties), structural_data, transaction_data, or relationships. In particular, datatype_properties will be organized separately in the next step. [Memo for relationships narrowed by the partition (if any)] If the partition result allows the range/domain of an existing relationship candidate to be narrowed (for example, the range of producesProduct can be refined from bdc:Product to the partition subclass bdc:FinishedGood), add that information as a free-form description memo to the relationships section of the YAML. Use the same tone (one-line description) as other candidates already persisted there. Formal definitions will come in the ontology design step. Example: relationships: - description: "WorkCenter → Product via routing chain (matched: 9,200 rows)" # JOIN verification candidate - description: "[partition memo] The range of producesProduct from WorkCenter can be narrowed to bdc:FinishedGood (Product partition result)" # partition memo If no new relationship memos arise from the partition, skip this step. ⚠️ After modifying the YAML, always confirm yaml.safe_load() parses successfully, and if it breaks, briefly explain what went wrong, fix it, and verify again."</CODE></PRE><PRE><CODE>Bad prompt: "Split the classes in this ontology appropriately." → "Appropriately" has no criterion. The LLM partitions by guesswork. → Information-carrying nodes such as BOMItem get oversimplified, losing information.</CODE></PRE><P class=""><STRONG>Roles:</STRONG></P><UL class=""><LI><STRONG>Domain expert</STRONG>: decides whether to partition, which column acts as the discriminator, and which class names to use</LI><LI><STRONG>LLM</STRONG>: identifying candidates from metadata (Step 1), persisting the decision as a YAML partition block (Step 2)</LI></UL><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":ledger:">📒</span><SPAN>&nbsp;</SPAN><STRONG>YAML changes from this step:</STRONG><SPAN>&nbsp;</SPAN>a<SPAN>&nbsp;</SPAN><CODE>partition</CODE><SPAN>&nbsp;</SPAN>block is added to the relevant DP entry. The existing<SPAN>&nbsp;</SPAN><CODE>kg_class</CODE><SPAN>&nbsp;</SPAN>(parent class, for example<SPAN>&nbsp;</SPAN><CODE>bdc:Product</CODE>) stays, and the discriminator column, the subclass mapping, and the disjoint flag are added. The Turtle class definitions themselves are generated when authoring ontology.ttl in the next step, reading this YAML.</P></BLOCKQUOTE><P class="">Section 4 deliverables:</P><P>&nbsp;</P><TABLE><TBODY><TR><TD><STRONG>Step</STRONG></TD><TD><STRONG>Deliverable</STRONG></TD><TD><STRONG>Role</STRONG></TD></TR><TR><TD>Class Partition decision</TD><TD><CODE>dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>update (discriminator partition rules + per-entity<SPAN>&nbsp;</SPAN><CODE>kg_class</CODE>)</TD><TD>Input for Section 6 ontology Turtle authoring</TD></TR></TBODY></TABLE><HR /><H1 id="5-completing-the-ontology" id="toc-hId-356105167">5. Completing the Ontology</H1><P class="">The KG ontology structure that emerges after applying every decision in Sections 3 and 4 is below. This structure is consistent with the final classification in<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE>.<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>becomes the<SPAN>&nbsp;</SPAN><CODE>bdc:MasterData</CODE><SPAN>&nbsp;</SPAN>hierarchy, PSL (the structural item with its own attributes) sits under<SPAN>&nbsp;</SPAN><CODE>bdc:StructuralNode</CODE>, and DPs that act only as intermediate bridges, such as ProductionRouting, have no KG nodes and serve only as KGR View inputs (Part 3). If the YAML was the working log, the ontology TTL is the result of translating that log into W3C standard syntax.</P><P class="">Instance counts below are SAP demo data examples.</P><PRE><CODE>rdfs:Class ├── bdc:MasterData │ ├── bdc:FinishedGood , finished good (225 examples) │ ├── bdc:SemiFinished , semi-finished (2,185 examples) │ ├── bdc:RawMaterial , raw material (183 examples) │ ├── bdc:Supplier , supplier (153 examples) │ ├── bdc:Plant , plant (12 examples) │ └── bdc:WorkCenter , work center (129 examples) │ └── bdc:StructuralNode └── bdc:PurchasingSourceList , Supplier ↔ Product bridge (16 examples)</CODE></PRE><P class="">Relationships:</P><PRE><CODE>WorkCenter ──bdc:locatedAt──▶ Plant WorkCenter ──bdc:producesProduct──▶ FinishedGood Product ──bdc:locatedAt──▶ Plant Product ──bdc:hasComponent──▶ Product PSL ──bdc:forProduct──▶ Product PSL ──bdc:fromSupplier──▶ Supplier</CODE></PRE><P class=""><CODE>PurchasingSourceList (PSL)</CODE><SPAN>&nbsp;</SPAN>is intentionally classified as<SPAN>&nbsp;</SPAN><CODE>StructuralNode</CODE><SPAN>&nbsp;</SPAN>rather than<SPAN>&nbsp;</SPAN><CODE>MasterData</CODE>. PSL is not a simple N:N connection; it carries information such as<SPAN>&nbsp;</SPAN><EM>"at which plant, for which validity period, with which priority."</EM><SPAN>&nbsp;</SPAN>The decision was to<SPAN>&nbsp;</SPAN><STRONG>keep meaningful junctions as graph nodes</STRONG><SPAN>&nbsp;</SPAN>rather than flattening them.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":pushpin:">📌</span><SPAN>&nbsp;</SPAN><STRONG>These relationships are realized as KGR Views in Part 3.</STRONG><SPAN>&nbsp;</SPAN><CODE>bdc:producesProduct</CODE>, which requires compressing a multi-hop chain;<SPAN>&nbsp;</SPAN><CODE>bdc:hasComponent</CODE>, which traverses a BOM;<SPAN>&nbsp;</SPAN><CODE>bdc:locatedAt</CODE>, a simple direct FK. Each is precomputed as a SQL View (<CODE>KGR_*</CODE>) that becomes input to instance TTL generation. The View patterns are covered in detail in Part 3 Section 3.</P></BLOCKQUOTE><H3 id="vibe-coding-organizing-relationships-in-the-yaml" id="toc-hId--427214352">Vibe Coding: organizing<SPAN>&nbsp;</SPAN><CODE>relationships</CODE><SPAN>&nbsp;</SPAN>in the YAML</H3><P class="">After Section 2 JOIN verification and Section 4 Class Partition, free-form description memos for candidates have accumulated in the<SPAN>&nbsp;</SPAN><CODE>relationships</CODE><SPAN>&nbsp;</SPAN>section of the YAML. In this step, those are organized in one pass into the<SPAN>&nbsp;</SPAN><STRONG>formal ObjectProperty form</STRONG><SPAN>&nbsp;</SPAN>(name / domain / range / source).</P><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE>: the<SPAN>&nbsp;</SPAN><CODE>relationships</CODE><SPAN>&nbsp;</SPAN>section with Section 2 JOIN verification candidates + Section 4 partition memos accumulated as free-form descriptions, plus the master_data / structural_data / partition regions</LI><LI>The Section 5 ontology diagram above (the confirmed relationship list — which relationships to keep or drop, and what to call them)</LI><LI>The verification results from<SPAN>&nbsp;</SPAN><CODE>ingestion/02_verify_joins.py</CODE><SPAN>&nbsp;</SPAN>(per-candidate JOIN pattern and matching row count) — the basis for authoring SQL View bodies</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>update: free-form items in<SPAN>&nbsp;</SPAN><CODE>relationships</CODE><SPAN>&nbsp;</SPAN>replaced with the formal form (name / domain / range / source)</LI><LI><CODE>ingestion/kgr_views/&lt;view_name lowercase&gt;.sql</CODE><SPAN>&nbsp;</SPAN>files — the View body for each relationship (only the SELECT portion of CREATE VIEW)</LI><LI>Change description (which free-form candidate became which formal relationship)</LI></UL></BLOCKQUOTE><PRE><CODE>Organize the relationships section of ingestion/dp_mapping.yaml. Free-form description candidates persisted in earlier steps (JOIN verification + Class Partition) have accumulated there. In this step, replace them in one pass with the formal ObjectProperty form. The ontology diagram I confirmed (relationship list): - WorkCenter → bdc:locatedAt → Plant - WorkCenter → bdc:producesProduct → bdc:FinishedGood (※ range narrowed by partition) - Product → bdc:locatedAt → Plant - Product → bdc:hasComponent → Product - PurchasingSourceList → bdc:forProduct → Product - PurchasingSourceList → bdc:fromSupplier → Supplier ... (adjust to your domain) Workflow: 1. Read every free-form description candidate in yaml relationships and match each to the confirmed list above. (Which candidate maps to which relationship.) 2. Replace relationships in the confirmed list with the formal form: - name : relationship name (camelCase) - domain : subject class - range : object class (use the narrowed subclass if a partition memo exists) - description : one-line description - source.type : kgr_view - source.view_name : convention KGR_&lt;SUBJ&gt;_&lt;REL&gt;_&lt;OBJ&gt; - source.sql_file : convention ingestion/kgr_views/&lt;view_name lowercase&gt;.sql 3. Remove free-form candidates that are not in the confirmed list. (Verified but decided against persisting in the ontology.) 4. For each formal relationship, author the actual SQL View body at the source.sql_file path: - Use the JOIN pattern from 02_verify_joins.py verification results - The file contains only the SELECT portion of CREATE VIEW (the 03 script wraps the CREATE VIEW statement) - Output columns: (subject_iri, object_iri) only. For multi-hop, route through intermediate JOINs and SELECT only the two final columns. Apply DISTINCT. - subject_iri / object_iri are IRI strings of the form 'bdc:&lt;Class&gt;_' || &lt;key column&gt; (for example, 'bdc:WorkCenter_' || wc.WorkCenterInternalID) ⚠️ Do NOT touch master_data / structural_data / transaction_data / partition / datatype_properties. They were decided in earlier steps or are scheduled for the next step. Organize the relationships section only. ⚠️ After modifying the YAML, always confirm yaml.safe_load() parses successfully, and if it breaks, briefly explain what went wrong, fix it, and verify again. Finally, show the result in a table: | free-form candidate (before) | formal relationship (after) | view_name | sql_file | notes |</CODE></PRE><P class=""><SPAN>The deliverables of this step come in pairs. In the YAML,&nbsp;</SPAN><CODE>relationships[*].source</CODE><SPAN>&nbsp;points to the view that realizes each relationship and to the&nbsp;</SPAN><CODE>.sql</CODE><SPAN>&nbsp;file that contains the view body. The actual SQL files are stored under&nbsp;</SPAN><CODE>ingestion/kgr_views/</CODE><SPAN>, one file per view. In the next post, Part 3, the KGR view creation script reads these YAML paths and loads each SQL body into HANA as a&nbsp;</SPAN><CODE>CREATE VIEW</CODE><SPAN>&nbsp;statement.</SPAN></P><H3 id="vibe-coding-organizing-datatype_properties-in-the-yaml" id="toc-hId--623727857">Vibe Coding: organizing<SPAN>&nbsp;</SPAN><CODE>datatype_properties</CODE><SPAN>&nbsp;</SPAN>in the YAML</H3><P class="">With the classes ready, the next task is to persist which columns will become attributes (<CODE>owl:DatatypeProperty</CODE>) on each class. The<SPAN>&nbsp;</SPAN><EM>"what to persist"</EM><SPAN>&nbsp;</SPAN>decision goes into the YAML; the actual Turtle definitions are produced by the LLM in Section 6 from the YAML.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE>: the six DPs in<SPAN>&nbsp;</SPAN><CODE>master_data</CODE><SPAN>&nbsp;</SPAN>(with relationships now populated)</LI><LI><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE>: column list + business meaning per DP</LI><LI>HANA<SPAN>&nbsp;</SPAN><CODE>PUBLIC.VIRTUAL_COLUMNS</CODE>:<SPAN>&nbsp;</SPAN><CODE>DATA_TYPE_NAME</CODE><SPAN>&nbsp;</SPAN>per column (for XSD type mapping)</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE><SPAN>&nbsp;</SPAN>update:<SPAN>&nbsp;</SPAN><CODE>master_data[*].datatype_properties</CODE><SPAN>&nbsp;</SPAN>populated for each DP (column / property / range)</LI><LI>Report of omitted / excluded columns (with reasons)</LI></UL></BLOCKQUOTE><PRE><CODE>Populate datatype_properties for each DP under master_data in ingestion/dp_mapping.yaml. [Input materials] - ingestion/dp_mapping.template.yaml : the datatype_properties form + XSD mapping rules are documented as comments. Follow that form exactly. - ingestion/dp_mapping.yaml : kg_class and partition for each DP under master_data are now decided - docs/dp_specs/CSN_SUMMARY.md : column list + business meaning per DP - Direct query of HANA PUBLIC.VIRTUAL_COLUMNS : XSD type mapping based on DATA_TYPE_NAME of rows whose SCHEMA_NAME is each DP's virtual_table_schema [XSD type mapping rules] - NVARCHAR / VARCHAR / STRING → xsd:string - INTEGER / BIGINT / SMALLINT / TINYINT → xsd:integer - DECIMAL / DOUBLE / REAL / FLOAT → xsd:decimal - DATE → xsd:date - TIMESTAMP / SECONDDATE → xsd:dateTime - BOOLEAN → xsd:boolean [Column selection rules] Include: - Name / Description (label · comment candidates) - business codes (ProductType, MaterialGroup, and so on) - amount / quantity / count - business dates (ValidFrom / ValidTo / DocumentDate, and so on) Exclude: - GUID / UUID (for example, ProductUUID, *_GUID) - audit metadata (CreatedAt / CreatedBy / LastChangedAt / LastChangedBy) - ETag / Version (for example, SAP_EntityChangeDate, SAP_ETag) - soft-delete flags (IsDeleted / IsArchived / DeletionIndicator) - FK columns (already persisted as ObjectProperties in relationships) [Form] Add a datatype_properties list to each DP under master_data: datatype_properties: - column: "&lt;source column name, e.g., WorkCenterName&gt;" property: "&lt;bdc:camelCase, e.g., bdc:workCenterName&gt;" range: "&lt;XSD type, e.g., xsd:string&gt;" ⚠️ Do NOT touch master_data / structural_data / transaction_data / partition / relationships. Add only master_data[*].datatype_properties. ⚠️ For DPs with a partition (for example, Product), persist datatype_properties at the parent class level (bdc:Product). Do not persist them separately per subclass (FinishedGood / SemiFinished / RawMaterial). [Self-verification] After modifying the YAML, confirm yaml.safe_load() parses. If it breaks, fix it and verify again. Finally, report: - How many columns were added per DP - The list of excluded columns + reason (GUID / audit / etc.)</CODE></PRE><P class="">A person scans the reported exclusion list and confirms only whether<SPAN>&nbsp;</SPAN><EM>"this column is actually domain-important and was dropped"</EM><SPAN>&nbsp;</SPAN>by mistake. In most cases, the rules apply cleanly.</P><P class="">The next Section 6 covers how this ontology structure is delegated to the LLM to produce the draft Turtle, and how labels and comments are written.</P><P class="">Section 5 deliverables:</P><TABLE><TBODY><TR><TD><STRONG>Step</STRONG></TD><TD><STRONG>Deliverable</STRONG></TD><TD><STRONG>Role</STRONG></TD></TR><TR><TD>Final ontology structure confirmed</TD><TD>Ontology diagram (6 MasterData + 1 StructuralNode + 6 relationships)</TD><TD>Blueprint for Section 6 Turtle authoring</TD></TR><TR><TD>relationships organized</TD><TD><CODE>dp_mapping.yaml</CODE>'s<SPAN>&nbsp;</SPAN><CODE>relationships</CODE><SPAN>&nbsp;</SPAN>section replaced from Section 2 / Section 4 free-form candidates with the formal form (name / domain / range / source)</TD><TD>Direct input for Section 6 Turtle authoring</TD></TR><TR><TD>KGR view SQL bodies authored</TD><TD><CODE>ingestion/kgr_views/&lt;view_name&gt;.sql</CODE><SPAN>&nbsp;</SPAN>files (one file per view, the SELECT body)</TD><TD>The Part 3 KGR view creation script reads them and loads them into HANA as CREATE VIEW</TD></TR><TR><TD>datatype_properties organized</TD><TD><CODE>dp_mapping.yaml</CODE>'s<SPAN>&nbsp;</SPAN><CODE>master_data[*].datatype_properties</CODE><SPAN>&nbsp;</SPAN>populated (column / property / range)</TD><TD>Direct input for Section 6 Turtle authoring</TD></TR></TBODY></TABLE><HR /><H1 id="6-drafting-the-ontology-turtle-with-ai-and-writing-labels" id="toc-hId--233435348">6. Drafting the Ontology Turtle with AI, and Writing Labels</H1><P class="">Working through Sections 2–5 accumulated every ontology decision in<SPAN>&nbsp;</SPAN><CODE>dp_mapping.yaml</CODE>: category classification, VT / synonym mapping,<SPAN>&nbsp;</SPAN><CODE>kg_class</CODE>,<SPAN>&nbsp;</SPAN><CODE>partition</CODE>, and<SPAN>&nbsp;</SPAN><CODE>relationships</CODE>. In this step, that YAML is converted to W3C standard Turtle format, while semantic vocabulary such as<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE>,<SPAN>&nbsp;</SPAN><CODE>rdfs:comment</CODE>, and<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>is added in parallel.</P><P class="">The key point is that<SPAN>&nbsp;</SPAN><STRONG>a person does not maintain the entity list in their head</STRONG>. The YAML is the SSoT of decisions, so a single LLM prompt,<SPAN>&nbsp;</SPAN><EM>"convert this YAML to Turtle"</EM><SPAN>,</SPAN>&nbsp;produces the draft.</P><BLOCKQUOTE dir="auto"><P class=""><STRONG><span class="lia-unicode-emoji" title=":inbox_tray:">📥</span>Input (entire Section 6)</STRONG></P><UL class=""><LI><CODE>ingestion/dp_mapping.yaml</CODE>: the SSoT accumulated across Sections 2–5 (category, kg_class, synonym, partition, relationships)</LI><LI><CODE>docs/dp_specs/CSN_SUMMARY.md</CODE>: columns, amounts, and quantities per DP — reference for the domain vocabulary in DatatypeProperty and label / comment</LI><LI>(Optional)<SPAN>&nbsp;</SPAN><CODE>02_verify_joins.py</CODE><SPAN>&nbsp;</SPAN>verification report: evidence that relationships are alive. Reference for memos such as row counts in comments</LI></UL><P class=""><STRONG><span class="lia-unicode-emoji" title=":outbox_tray:">📤</span>Output (entire Section 6)</STRONG></P><UL class=""><LI><CODE>ingestion/04_generate_ontology.py</CODE>: a Python script that persists the TTL text as a multi-line string and writes it to<SPAN>&nbsp;</SPAN><CODE>ingestion/ontology/bdc_dp_ontology.ttl</CODE><SPAN>&nbsp;</SPAN>(using only<SPAN>&nbsp;</SPAN><CODE>pathlib</CODE>, no external libraries such as rdflib)</LI><LI><CODE>ingestion/ontology/bdc_dp_ontology.ttl</CODE>: the script output — classes + ObjectProperty + DatatypeProperty + multilingual labels and comments + skos:altLabel synonyms</LI><LI>Used in: Part 3 Section 5 TTL generation (loaded into HANA KGE)</LI></UL></BLOCKQUOTE><H3 id="vibe-coding-yaml--ontologyttl-generation" id="toc-hId--1016754867">Vibe Coding: YAML → ontology.ttl generation</H3><PRE><CODE>Write a Python script ingestion/04_generate_ontology.py that takes ingestion/dp_mapping.yaml as input and generates a W3C Turtle ontology. [Output form] - Persist the entire TTL inside the Python script as a multi-line string (ONTOLOGY = \"\"\"...\"\"\") and write it to ingestion/ontology/bdc_dp_ontology.ttl. - Use only pathlib, no external libraries. Do NOT use rdflib for builder code. Author the TTL text yourself and persist it as a string. [Input materials] - ingestion/dp_mapping.yaml : the SSoT of every ontology decision - master_data[*].kg_class : class name - master_data[*].partition : Class Partition (generate subclasses if present) - master_data[*].datatype_properties : DatatypeProperty definitions (column, property, range) - structural_data[*].kg_class : StructuralNode hierarchy - relationships[*] : ObjectProperty definitions (name, domain, range, description) - docs/dp_specs/CSN_SUMMARY.md : reference for label / comment domain vocabulary (the names and types of datatype_properties items in the YAML are already decided; only the label / comment phrasing in business vocabulary comes from here) Authoring rules: 1. Prefix declarations (bdc:, rdfs:, rdf:, owl:, skos:, xsd:) 2. Class hierarchy: - kg_class values in master_data are rdfs:subClassOf bdc:MasterData - kg_class values in structural_data are subclasses of bdc:StructuralNode - DPs with a partition have the subclasses under partition.parent_class as rdfs:subClassOf; if partition.disjoint=true, also owl:disjointWith 3. Add rdfs:label and rdfs:comment to every class (multilingual: @en, @ko) 4. Define each item in relationships as owl:ObjectProperty (domain / range / label / comment) 5. Define each item in master_data[*].datatype_properties as owl:DatatypeProperty: - bdc:&lt;property&gt; a owl:DatatypeProperty ; rdfs:domain bdc:&lt;kg_class&gt; ; rdfs:range &lt;range&gt; ; rdfs:label ... ; rdfs:comment ... . - Take label / comment from the column descriptions in CSN_SUMMARY (multilingual @en/@ko). ⚠️ Do NOT add skos:altLabel (synonyms) in this step. Handled separately. ⚠️ Follow the YAML exactly. Do not add relationships or classes freely. ⚠️ Do NOT persist implementation metadata (view_name / sql_file) from relationships[*].source into the Turtle. The ontology contains semantics only (domain / range / label / comment). Implementation assets (views, SQL) are managed in separate .sql files. [Self-verification — not inside the script, but as your own self-check] Run the script once and parse the generated TTL with rdflib.Graph().parse() to confirm it parses. If it breaks, briefly explain what went wrong, fix it, and verify again. Hand over the final version in a parsing-passed state.</CODE></PRE><P class=""><STRONG>Roles:</STRONG></P><UL class=""><LI><STRONG>LLM</STRONG>: authoring Turtle from the YAML, maintaining label · comment · domain · range consistency, multilingual expression</LI><LI><STRONG>Domain expert</STRONG>: skims the class hierarchy and relationship direction in a visualization tool to confirm they match the YAML decisions, and reviews the business tone and multilingual phrasing of labels and comments</LI></UL><H3 id="reviewing-the-ontology-with-visualization--modeling-tools" id="toc-hId--1213268372">Reviewing the ontology with visualization / modeling tools</H3><P class="">A person reviews the tone and multilingual phrasing of labels and comments. The class hierarchy, ObjectProperty definitions, and domain / range constraints were decided in the YAML, so the LLM's output can be trusted as-is.</P><P class="">Before loading the draft from the LLM, a visual review is a safety net. Confirm graphically that the class hierarchy, ObjectProperty direction, and domain / range constraints came out as intended. A few options:</P><UL class=""><LI><STRONG>Protégé</STRONG>: open-source ontology editor developed at Stanford. Opens a TTL file and shows the class hierarchy and ObjectProperty as a tree. Useful for a light review.</LI><LI><STRONG>SAP HANA Cloud Knowledge Graph Engine</STRONG>: query the loaded KG directly with SPARQL for verification. Because ontology and instances live on a single platform, post-load verification is fast.</LI><LI><STRONG><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/sap-hana-cloud-integration-with-metaphactory-for-knowledge-graph-modelling/ba-p/14390776" target="_blank">Metaphactory</A></STRONG>: an Enterprise KG platform officially integrated with SAP HANA Cloud KGE. Supports ontology editing, visual modeling, data entry, and SPARQL queries in a single interface. Suited for operating an ontology in a collaborative environment.</LI></UL><P class=""><STRONG>In a production project</STRONG>, ontology design and long-term maintenance should be done with modeling and governance tools such as Protégé or Metaphactory. The AI-based workflow in this post is not intended to replace those tools. Its value is that it creates a strong first draft quickly, especially when you are exploring a new domain or learning how SAP BDC metadata maps into KG classes and relationships. For scenario-based projects, this first draft is often sufficient to validate the core classes, relationships, and Competency Questions before formalizing the ontology in dedicated modeling tools. Once you gain experience with this workflow, you will be better prepared to review, refine, and operate the ontology in those tools. This is also aligned with where the tooling market is moving: modeling environments increasingly use AI to assist semantic modeling, vocabulary creation, and relationship discovery.</P><P class=""><STRONG>In our project</STRONG>, we reviewed with Protégé, then generated the final TTL with<SPAN>&nbsp;</SPAN><CODE>04_generate_ontology.py</CODE><SPAN>&nbsp;</SPAN>(covered in detail in the next post, Part 3).</P><H3 id="synonyms-applying-skosaltlabel-in-practice" id="toc-hId--1241598186">Synonyms: applying skos:altLabel in practice</H3><P class="">Part 2-1 Section 1 covered the standards basis of the<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>pattern. Here we examine how it is applied in practice.</P><P class=""><STRONG>Situations where synonyms are needed:</STRONG></P><UL class=""><LI>When the Agent must link diverse expressions in user questions —<SPAN>&nbsp;</SPAN><EM>"공급사,"</EM><SPAN>&nbsp;</SPAN><EM>"벤더,"</EM><SPAN>&nbsp;</SPAN><EM>"Vendor,"</EM><SPAN>&nbsp;</SPAN><EM>"Supplier"</EM><SPAN>&nbsp;</SPAN>— to the same class</LI><LI>When<SPAN>&nbsp;</SPAN><CODE>context_loader.py</CODE><SPAN>&nbsp;</SPAN>injects ontology metadata into the LLM, so that the LLM generates the correct SPARQL regardless of the user's vocabulary</LI></UL><PRE><CODE>@prefix skos: &lt;http://www.w3.org/2004/02/skos/core#&gt; . bdc:Supplier a owl:Class ; rdfs:label "Supplier"@en , "공급사"@ko ; skos:altLabel "공급업체"@ko , "벤더"@ko , "Vendor"@en ; rdfs:comment "A counterparty that supplies materials or services."@en .</CODE></PRE><P class="">This is the step deliberately omitted from the first Vibe Coding. Synonyms depend on industry, domain, and in-house vocabulary, so handing them off in a separate pass after the ontology draft is in place produces better quality.</P><BLOCKQUOTE dir="auto"><P class=""><span class="lia-unicode-emoji" title=":direct_hit:">🎯</span><SPAN>&nbsp;</SPAN><STRONG>Vibe Coding: bulk addition of skos:altLabel synonyms</STRONG></P><P class=""><EM>Purpose:</EM><SPAN>&nbsp;</SPAN>add user vocabulary variants (synonyms) to the ontology draft from the first Vibe Coding as<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE></P><P class=""><EM>Input:</EM></P><UL class=""><LI><CODE>ingestion/04_generate_ontology.py</CODE><SPAN>&nbsp;</SPAN>(the first Vibe Coding deliverable — the TTL string is persisted as a multi-line string)</LI><LI>Domain context (SAP industry vocabulary / in-house jargon / English vs Korean phrasing differences)</LI></UL><P class=""><EM>Deliverable:</EM><SPAN>&nbsp;</SPAN>updated<SPAN>&nbsp;</SPAN><CODE>04_generate_ontology.py</CODE><SPAN>&nbsp;</SPAN>—<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>triples added to the TTL string and<SPAN>&nbsp;</SPAN><CODE>@prefix skos:</CODE><SPAN>&nbsp;</SPAN>added to the top prefix block. Re-running updates<SPAN>&nbsp;</SPAN><CODE>bdc_dp_ontology.ttl</CODE>.</P></BLOCKQUOTE><PRE><CODE>Edit the TTL string (ONTOLOGY = \"\"\"...\"\"\") inside ingestion/04_generate_ontology.py directly to add skos:altLabel synonyms to each class. The targets are every class in master_data and structural_data plus the subclasses of every Class Partition. Rules: 1. Add @prefix skos: &lt;http://www.w3.org/2004/02/skos/core#&gt; . to the top prefix block 2. Leave rdfs:label as the "official name"; add only the other expressions users actually use as skos:altLabel 3. If both Korean and English variants are possible, add both (keep @ko / @en language tags) 4. SAP industry vocabulary (for example, Supplier → Vendor/벤더/거래처, Plant → 공장/사업장, BillOfMaterial → BOM/명세서) 5. If you are uncertain for a given class, skip it and report which classes were skipped at the end 6. Do NOT introduce external libraries such as rdflib. Edit the existing multi-line string directly. [Self-verification] After re-running the script, parse the generated TTL once with rdflib.Graph().parse() to confirm. If it breaks, fix and verify again. Hand over the final version.</CODE></PRE><P class="">A person reviews only whether industry-specific terms or in-house jargon were omitted. If the LLM reports any<SPAN>&nbsp;</SPAN><EM>"skipped classes,"</EM><SPAN>&nbsp;</SPAN>fill in the in-house vocabulary directly and run one more pass.</P><P class=""><STRONG>Synonyms are possible without SKOS.</STRONG><SPAN>&nbsp;</SPAN>A custom property, or packing every variant into<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE>, also makes search work. The reason to use<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>is not to enable search; it is two-fold: the preferred / alternative role distinction is standardized and built in, and when later extending into cross-system mapping (<CODE>skos:exactMatch</CODE><SPAN>&nbsp;</SPAN>and similar), the same vocabulary family is reused. For the detailed trade-off, consult<SPAN>&nbsp;</SPAN><STRONG>Part 2-1 Section 1, "Can synonyms work without SKOS?"</STRONG></P><H3 id="why-this-work-is-done-at-the-ontology-stage-it-is-reused-in-the-agent" id="toc-hId--1438111691">Why this work is done at the ontology stage: it is reused in the Agent</H3><P class="">Labels and comments are not just readability aids.<SPAN>&nbsp;</SPAN><STRONG>They are the input that is injected as domain vocabulary into the LLM in the Part 4 Agent</STRONG>:</P><UL class=""><LI><STRONG>Ontology metadata loaded at Agent startup</STRONG>: class labels and comments are persisted in the system prompt so the LLM learns<SPAN>&nbsp;</SPAN><EM>"<CODE>bdc:WorkCenter</CODE><SPAN>&nbsp;</SPAN>is a work center located in a Plant, producing Finished Goods"</EM><SPAN>&nbsp;</SPAN>→ smarter SPARQL authoring</LI><LI><STRONG>SPARQL result interpretation</STRONG>: with only URIs (<CODE>bdc:WorkCenter_Z_ASM3</CODE>), the LLM cannot read; with a label<SPAN>&nbsp;</SPAN><EM>"Z_ASM3"</EM>, it generates an answer using the business name directly</LI><LI><STRONG>Multilingual answers</STRONG>: a Korean question prompts the Agent to retrieve the<SPAN>&nbsp;</SPAN><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/137618">@KO</a></CODE><SPAN>&nbsp;</SPAN>label / comment first and answer in Korean, and English for English speakers. Multilingual support is delivered through the ontology alone, without backend code changes.</LI></UL><P class="">→ The quality of labels and comments at this stage is the quality of the Agent's answers.<SPAN>&nbsp;</SPAN><EM>"Polish them later"</EM><SPAN>&nbsp;</SPAN>is not the right approach; persist them alongside the ontology load. That is the standard.</P><P class="">Section 6 deliverables:</P><TABLE><TBODY><TR><TD><STRONG>Step</STRONG></TD><TD><STRONG>Deliverable</STRONG></TD><TD><STRONG>Role</STRONG></TD></TR><TR><TD>Ontology draft</TD><TD><CODE>ingestion/ontology/bdc_dp_ontology.ttl</CODE><SPAN>&nbsp;</SPAN>(draft)</TD><TD>Classes + ObjectProperty + DatatypeProperty definitions</TD></TR><TR><TD>Multilingual label · comment authoring</TD><TD><CODE>bdc_dp_ontology.ttl</CODE><SPAN>&nbsp;</SPAN>(final)</TD><TD><CODE>@en</CODE><SPAN>&nbsp;</SPAN>/<SPAN>&nbsp;</SPAN><CODE>@ko</CODE><SPAN>&nbsp;</SPAN>multilingual labels and comments +<SPAN>&nbsp;</SPAN><CODE>skos:altLabel</CODE><SPAN>&nbsp;</SPAN>synonyms</TD></TR></TBODY></TABLE><P class="">→ In the next post (Part 3), this<SPAN>&nbsp;</SPAN><CODE>bdc_dp_ontology.ttl</CODE><SPAN>&nbsp;</SPAN>is loaded into SAP HANA Cloud KGE.</P><HR /><H2 id="summary" id="toc-hId--1341222189">Summary</H2><P class="">This post covered:</P><UL class=""><LI><STRONG>KG vs SQL classification</STRONG>: structural relationship traversal → KG, numeric aggregation → SQL. Entities with no outgoing relationships (Customer) become SQL lookups</LI><LI><STRONG>Class Partition</STRONG>: separating different kinds of entities within a single table (Product by MaterialType) into distinct classes. Relationship directionality + LLM reasoning</LI><LI><STRONG>Final ontology</STRONG>: two branches (MasterData / StructuralNode), six master classes + one bridge (PSL)</LI><LI><STRONG>Drafting with AI</STRONG>: SAP ODM vocabulary lets the LLM produce a strong draft</LI><LI><STRONG><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>authoring rules</STRONG>: applied to every instance, class, and property. The minimum requirement that makes the KG readable by the LLM</LI></UL><P class="">The deliverable of this post is a single file,<SPAN>&nbsp;</SPAN><CODE>bdc_dp_ontology.ttl</CODE><SPAN>&nbsp;,</SPAN>&nbsp;the blueprint of the KG. The next post covers the pipeline that fills that blueprint with actual data.</P><HR /><H2 id="whats-next" id="toc-hId--1537735694">What's Next</H2><P class="">The blueprint is complete. What remains is reading actual data from SAP BDC Data Products, transforming it into the KG, and loading it into SAP HANA Cloud KGE.</P><P class=""><STRONG>Part 3 – Data Pipeline from BDC to Triple Store</STRONG><SPAN>&nbsp;</SPAN>starts with metadata exploration using<SPAN>&nbsp;</SPAN><CODE>PUBLIC.VIRTUAL_COLUMNS</CODE>, covers the relationship SQL View pattern that compresses multi-hop relationships into 1-hop, the TTL generation pipeline, and the Triple Store inside SAP HANA Cloud KGE.</P><HR /><P class=""><EM>Stack: SAP BDC Data Products · HANA Cloud KGE · FastAPI · React · Claude (Anthropic / AI Core)</EM><SPAN>&nbsp;</SPAN><EM>Code:<SPAN>&nbsp;</SPAN><A href="https://github.com/claudiopark86/sap-bdc-kge-agent-workshop" target="_blank" rel="noopener nofollow noreferrer">github.com/claudiopark86/sap-bdc-kge-agent-workshop</A></EM></P><HR /><H2 id="references" id="toc-hId--1734249199">References</H2><UL class=""><LI><STRONG>Ontology Development 101: A Guide to Creating Your First Ontology</STRONG><SPAN>&nbsp;</SPAN>— Noy &amp; McGuinness, Stanford 2001.<SPAN>&nbsp;</SPAN><A href="https://protege.stanford.edu/publications/ontology_development/ontology101.pdf" target="_blank" rel="noopener nofollow noreferrer">https://protege.stanford.edu/publications/ontology_development/ontology101.pdf</A></LI><LI><STRONG>A Direct Mapping of Relational Data to RDF</STRONG><SPAN>&nbsp;</SPAN>— W3C Recommendation, 2012-09-27.<SPAN>&nbsp;</SPAN><A href="https://www.w3.org/TR/rdb-direct-mapping/" target="_blank" rel="noopener nofollow noreferrer">https://www.w3.org/TR/rdb-direct-mapping/</A></LI><LI><STRONG>R2RML: RDB to RDF Mapping Language</STRONG><SPAN>&nbsp;</SPAN>— W3C Recommendation, 2012-09-27.<SPAN>&nbsp;</SPAN><A href="https://www.w3.org/TR/r2rml/" target="_blank" rel="noopener nofollow noreferrer">https://www.w3.org/TR/r2rml/</A></LI><LI><STRONG>OWL 2 Quick Reference Guide</STRONG><SPAN>&nbsp;</SPAN>— W3C Recommendation, 2012-12-11. (§2.7 Annotations lists<SPAN>&nbsp;</SPAN><CODE>rdfs:label</CODE><SPAN>&nbsp;</SPAN>and<SPAN>&nbsp;</SPAN><CODE>rdfs:comment</CODE><SPAN>&nbsp;</SPAN>as built-in annotation properties.)<SPAN>&nbsp;</SPAN><A href="https://www.w3.org/TR/owl2-quick-reference/" target="_blank" rel="noopener nofollow noreferrer">https://www.w3.org/TR/owl2-quick-reference/</A></LI><LI><STRONG>SAP HANA Cloud Integration with Metaphactory for Knowledge Graph Modelling</STRONG><SPAN>&nbsp;</SPAN>— SAP Community Blog.<SPAN>&nbsp;</SPAN><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/sap-hana-cloud-integration-with-metaphactory-for-knowledge-graph-modelling/ba-p/14390776" target="_blank">https://community.sap.com/t5/technology-blog-posts-by-sap/sap-hana-cloud-integration-with-metaphactory-for-knowledge-graph-modelling/ba-p/14390776</A></LI><LI><STRONG>HANA Cloud Data Products Consumption</STRONG><SPAN>&nbsp;</SPAN>— SAP Tutorial.<SPAN>&nbsp;</SPAN><A href="https://developers.sap.com/tutorials/hana-cloud-data-products-consumption.html" target="_blank" rel="noopener noreferrer">https://developers.sap.com/tutorials/hana-cloud-data-products-consumption.html</A></LI><LI><STRONG>Data Product Support in SAP HANA Cloud</STRONG><SPAN>&nbsp;</SPAN>— SAP Help Portal.<SPAN>&nbsp;</SPAN><A href="https://help.sap.com/docs/hana-cloud/sap-hana-cloud-administration-guide/data-product-support-in-sap-hana-cloud-internal?locale=en-US" target="_blank" rel="noopener noreferrer">https://help.sap.com/docs/hana-cloud/sap-hana-cloud-administration-guide/data-product-support-in-sap-hana-cloud-internal?locale=en-US</A></LI></UL> 2026-07-04T15:45:25.498000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/code-connect-2026-final-info-before-we-kick-off/ba-p/14434082 Code Connect 2026: Final Info Before We Kick Off! 2026-07-06T11:02:20.615000+02:00 BirgitS https://community.sap.com/t5/user/viewprofilepage/user-id/41902 <P><SPAN>Code Connect 2026 is just around the corner! From <STRONG>July 13–16, 2026</STRONG>, the SAP developer community will gather in <STRONG>St. Leon-Rot, Germany</STRONG> – onsite and online.</SPAN></P><P><SPAN>Here’s everything you need to know before we get started.</SPAN></P><P>&nbsp;</P><P><SPAN><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Code Connect 2026" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429665i8288C16ACE3EC08A/image-size/large?v=v2&amp;px=999" role="button" title="BirgitS_0-1783324399225.png" alt="BirgitS_0-1783324399225.png" /></span></SPAN></P><P>&nbsp;</P><H2 id="toc-hId-1819132728">We Are Fully Booked!</H2><P><SPAN>All onsite tickets are <STRONG>fully booked</STRONG> - thank you for the amazing interest!</SPAN></P><P><SPAN>If you couldn’t secure a ticket:</SPAN></P><UL><LI><SPAN>Join the <STRONG>waiting lists on </STRONG><A href="https://code-connect.dev/" target="_blank" rel="noopener nofollow noreferrer"><STRONG>Code Connect 2026</STRONG></A> in case seats become available.</SPAN></LI><LI><SPAN>Follow selected sessions via <STRONG>live stream</STRONG> or <STRONG>Microsoft Teams</STRONG> link.</SPAN></LI></UL><P><SPAN>Even if you're not onsite, you can still be part of Code Connect. Selected sessions will be streamed live, and recordings of some sessions will be available afterwards. Check back soon for details.</SPAN></P><P>&nbsp;</P><TABLE><TBODY><TR><TD width="200px"><P><span class="lia-inline-image-display-wrapper lia-image-align-left" image-alt="UI5con" style="width: 76px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429685i9BABBF665F151967/image-size/small?v=v2&amp;px=200" role="button" title="UI5conSmall.png" alt="UI5conSmall.png" /></span></P><P>&nbsp;</P><P>&nbsp;</P><P><STRONG><SPAN>UI5con (July 14)</SPAN></STRONG></P></TD><TD width="465.047px"><UL><LI><SPAN>Main stage sessions will be live streamed on <A href="https://www.youtube.com/live/CMPudw4scSE" target="_blank" rel="noopener nofollow noreferrer">YouTube</A>. </SPAN></LI><LI><SPAN>Check the <A href="https://openui5.org/ui5con/program.html" target="_blank" rel="noopener nofollow noreferrer">agenda</A> to see which sessions are streamed.</SPAN></LI></UL></TD></TR><TR><TD width="200px"><P><STRONG><span class="lia-inline-image-display-wrapper lia-image-align-left" image-alt="reCAP" style="width: 76px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429686i805B9E02FF474F62/image-size/small?v=v2&amp;px=200" role="button" title="reCAPSmall.png" alt="reCAPSmall.png" /></span></STRONG></P><P>&nbsp;</P><P>&nbsp;</P><P><STRONG>re&gt;≡CAP (July 15)</STRONG></P></TD><TD width="465.047px"><UL><LI><SPAN>Main stage sessions will be <A href="https://broadcast.sap.com/go/reCAP" target="_blank" rel="noopener noreferrer">broadcasted</A>. </SPAN></LI><LI><SPAN>Some sidetracks (W1/W2) can be attended via <A href="https://teams.microsoft.com/meet/347149883889697?p=ZkR4erzGjpPRtZ6zJh" target="_blank" rel="noopener nofollow noreferrer">Microsoft Teams link</A>. </SPAN></LI><LI><SPAN>Check the <A href="https://recap-conf.dev/program.html" target="_blank" rel="noopener nofollow noreferrer">agenda</A> to see which sessions are broadcasted or available via Microsoft Teams link.</SPAN></LI></UL></TD></TR><TR><TD width="200px"><P><STRONG><SPAN><span class="lia-inline-image-display-wrapper lia-image-align-left" image-alt="HANATechCon" style="width: 76px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429687iDED26458B98CF5BC/image-size/small?v=v2&amp;px=200" role="button" title="HANATechConSmall.png" alt="HANATechConSmall.png" /></span></SPAN></STRONG></P><P>&nbsp;</P><P>&nbsp;</P><P><STRONG><SPAN>HANA Tech Con (July 16)</SPAN></STRONG></P></TD><TD width="465.047px"><UL><LI><SPAN>The links will be announced on the event day.</SPAN></LI></UL></TD></TR></TBODY></TABLE><P>&nbsp;</P><H2 id="toc-hId-1622619223">The Agendas Are Live</H2><P><SPAN>The full program is already available - time to plan your schedule!</SPAN></P><P><SPAN>Check-In opens at 8.00 AM every day. Please bring the QR code with you that you can find in your ticket.</SPAN></P><P><STRONG><SPAN>July 13 – Code Jams &amp; Community Meetup</SPAN></STRONG></P><UL><LI><SPAN>Hands-on <STRONG>Code Jams</STRONG>:</SPAN></LI><UL><LI><STRONG><SPAN>OpenUI5</SPAN></STRONG></LI><LI><STRONG><SPAN>CAP</SPAN></STRONG></LI><LI><STRONG><SPAN>AI Agents</SPAN></STRONG></LI></UL><LI><SPAN>Informal <STRONG>pre-event meetup</STRONG> (no registration needed):<BR />Ihle Besen, Höfe am Sträßel 3, 69231 Rauenberg<BR />5:00 PM CEST</SPAN></LI></UL><P><STRONG><SPAN>July 14–16 – Main Conference Days</SPAN></STRONG></P><P><SPAN>Explore all agendas:</SPAN></P><UL><LI><STRONG><SPAN>July 14:</SPAN></STRONG><SPAN> <A href="https://openui5.org/ui5con/program.html" target="_blank" rel="noopener nofollow noreferrer">UI5con</A></SPAN></LI><LI><STRONG><SPAN>July 15:</SPAN></STRONG><SPAN> <A href="https://recap-conf.dev/program.html" target="_blank" rel="noopener nofollow noreferrer">re&gt;≡CAP</A> </SPAN></LI><LI><STRONG><SPAN>July 16:</SPAN></STRONG><SPAN> <A href="https://hanatech.community/" target="_blank" rel="noopener nofollow noreferrer">HANA Tech Con</A></SPAN></LI></UL><P><SPAN>On <STRONG>July 14</STRONG> also the <STRONG>HANA AI CodeJam</STRONG> takes place. </SPAN></P><P><SPAN>&nbsp;</SPAN></P><H2 id="toc-hId-1426105718">Location</H2><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Location" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/429667i854A5327BE35386B/image-size/large?v=v2&amp;px=999" role="button" title="BirgitS_4-1783324399247.png" alt="BirgitS_4-1783324399247.png" /></span></P><P><SPAN>Code Connect takes place at:</SPAN></P><P><SPAN>SAP-Allee 27<BR />68789 St. Leon-Rot, Germany</SPAN></P><P><SPAN>Find detailed information on <STRONG>directions, parking, and public transport</STRONG> <A href="https://code-connect.dev/location.html" target="_blank" rel="noopener nofollow noreferrer">here.</A></SPAN></P><P>&nbsp;</P><H2 id="toc-hId-1229592213"><SPAN>Final Tips</SPAN></H2><UL><LI>Bookmark your favorite sessions in advance</LI><LI><SPAN>Arrive early for popular sessions (rooms fill quickly!)</SPAN></LI><LI>Bring your laptop for Code Jams and hands-on sessions/ workshops</LI><LI><SPAN>Don’t miss the networking moments - some of the best conversations happen in between sessions</SPAN></LI><LI><SPAN>Follow the agendas closely for last-minute updates</SPAN></LI></UL><P><SPAN>&nbsp;</SPAN></P><H2 id="toc-hId-1033078708">See You Soon!</H2><P><SPAN>We’re looking forward to an inspiring week of learning, sharing, and connecting.</SPAN></P><P><STRONG><SPAN>See you at Code Connect 2026!</SPAN></STRONG></P> 2026-07-06T11:02:20.615000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/building-vector-search-applications-with-sap-hana-cloud-and-langchain-js/ba-p/14413474 Building Vector Search Applications with SAP HANA Cloud and LangChain.js 2026-07-07T11:27:33.345000+02:00 varin_thakur https://community.sap.com/t5/user/viewprofilepage/user-id/2247820 <P>Imagine your customer types <EM>“something comfortable for working from home”</EM> into your shop’s search bar. There’s no “comfortable” column in your product table. No “work from home” tag either. And yet your store somehow needs to surface the ergonomic chair, the noise-cancelling headphones, and the standing desk.</P><P>That gap between <STRONG>what people say</STRONG> and <STRONG>what your database stores</STRONG> is exactly where vector search lives. With <STRONG><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/2302137">@SAP</a>/hana-langchain</CODE></STRONG>, your SAP HANA Cloud database doesn’t just store your data. It understands it.</P><BLOCKQUOTE><P><STRONG>The full walkthrough, with every line of code, every output, and every config option, lives on GitHub:</STRONG> <A href="https://github.com/SAP/langchainjs-integration-for-sap-hana-cloud/blob/main/blogs/blog-post.md" target="_blank" rel="noopener nofollow noreferrer">blogs/blog-post.md</A>. The sections below focus on the ideas worth remembering, paired with the bits of code and output that bring them to life.</P></BLOCKQUOTE><H2 id="toc-hId-1817259710">What it is</H2><P><CODE><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/2302137">@SAP</a>/hana-langchain</CODE> is the official open-source bridge between <STRONG>SAP HANA Cloud’s Vector Engine</STRONG> and <STRONG>LangChain.js</STRONG>. It turns your enterprise database into a first-class AI retrieval layer for RAG, semantic search, recommendations, and intelligent Q&amp;A, all running where your data already lives.</P><P>Two lines and you have a vector store:</P><PRE><CODE>npm install @langchain/core @langchain/classic langchain npm install @sap/hana-langchain</CODE></PRE><H2 id="toc-hId-1620746205">Searching by intent, not by keywords</H2><P>Let’s walk through the real story. A customer is setting up their home office. They open your store and search:</P><BLOCKQUOTE><P><EM>“comfortable work from home setup”</EM></P></BLOCKQUOTE><P>Your catalogue has hundreds of products. Keyword search? Zero relevant hits. With <CODE>HanaDB</CODE>? Watch what happens.</P><PRE><CODE>import { HanaDB, HanaInternalEmbeddings } from "@sap/hana-langchain"; // Embeddings happen INSIDE HANA. No external API. No data leaving your DB. const embeddings = new HanaInternalEmbeddings({ internalEmbeddingModelId: "SAP_NEB.20240715", }); const store = new HanaDB(embeddings, { connection: client, tableName: "PRODUCT_CATALOG", }); await store.initialize(); const results = await store.similaritySearch( "comfortable work from home setup", 2, );</CODE></PRE><P>The output:</P><PRE><CODE>[ Document { pageContent: 'Ergonomic office chair with lumbar support and adjustable armrests', metadata: { category: 'furniture', price: 449, in_stock: true } }, Document { pageContent: 'Professional espresso machine with built-in grinder', metadata: { category: 'appliances', price: 899.99, in_stock: false } } ]</CODE></PRE><P>The chair makes obvious sense. The espresso machine? Also surprisingly on-brand for “working from home.” The model picked up on intent that no SQL <CODE>LIKE</CODE> ever could.</P><H2 id="toc-hId-1424232700">Cosine vs. Euclidean: pick your distance</H2><P>So how does <CODE>similaritySearch</CODE> actually decide which products are “closest” to the query? Each document and the query are turned into vectors, and then <CODE>HanaDB</CODE> applies a <EM>distance metric</EM> to measure how far apart they are. The library supports two: <STRONG>cosine</STRONG> (the default) and <STRONG>Euclidean</STRONG>. You pick which one you want with a single config option when you create the store.</P><P>The choice matters because the two metrics score the same set of candidates differently. Take the query <EM>“high quality audio headphones”</EM> and run it under both strategies. Same data, same embedding model, different scoring philosophy:</P><PRE><CODE>COSINE Similarity Results (higher = more similar) [0.7219] Wireless noise-canceling headphones with premium sound quality [0.5506] Professional studio microphone for podcasting and streaming [0.5206] Bluetooth speaker with deep bass and 20-hour battery EUCLIDEAN Distance Results (lower = more similar) [0.7086] Wireless noise-canceling headphones with premium sound quality [0.9048] Professional studio microphone for podcasting and streaming [0.9251] Bluetooth speaker with deep bass and 20-hour battery</CODE></PRE><P>Both metrics pick the same winner here, but notice the gap between first and second place: cosine spreads the candidates by relative similarity, Euclidean by raw distance in vector space. As a rule of thumb, <STRONG>cosine</STRONG> is the right default for text, because it only cares about the <EM>direction</EM> the vectors point in, not how long they are. <STRONG>Euclidean</STRONG> is a better fit when the actual position of the vectors carries meaning, for example when your embeddings come from numerical features rather than text. If you’re not sure which to pick, leave it on cosine.</P><H2 id="toc-hId-1227719195">Adding diversity when you need it: MMR</H2><P>Sometimes “most similar” is exactly what the user wants. A shopper who searches for <EM>“wireless headphones”</EM> expects a list of wireless headphones, and similarity search delivers that perfectly well.</P><P>But there are spots in the same product journey where pure similarity hurts the experience. The classic case is the <STRONG>“customers also viewed”</STRONG> strip on a product detail page. A shopper looking at a premium open-back studio headphone doesn’t need to see four more premium over-ear headphones below it. They’d be much better served by a strip that <EM>also</EM> shows them a different form factor, a different price tier, or a noise-cancelling alternative.</P><P>That’s the job of <STRONG>Maximal Marginal Relevance</STRONG> (MMR). Instead of returning the top-k most similar documents, MMR picks results that are <EM>both</EM> relevant <EM>and</EM> different from each other. You retrieve a larger candidate pool; MMR walks it, picking the next item that maximises relevance while penalising redundancy with what’s already been chosen. One <CODE>lambda</CODE> knob tunes the balance between relevance and diversity.</P><P>Same vector store, same query, two different methods. First, plain similarity search:</P><PRE><CODE>// Plain similarity search: top-4 candidates ordered by closeness alone const similar = await vectorStore.similaritySearch( "premium wireless over-ear headphones", 4, );</CODE></PRE><PRE><CODE>similaritySearch top 4: 1. AudioMax Pro 5 (over-ear-wireless, $349.99) 2. AuroraSound Pro (over-ear-wireless, $379.99) 3. AudioMax Buds 5 (true-wireless, $299.99) 4. QuietShield Ultra (over-ear-wireless, $429.99)</CODE></PRE><P>Three of the four results are premium over-ear wireless headphones, almost interchangeable for the shopper. Not a useful “customers also viewed” strip.</P><P>Now the same query, through MMR:</P><PRE><CODE>// MMR: pick 4 from a wider candidate pool while penalising redundancy. // lambda = 0.5 balances relevance and diversity evenly. const recommendations = await vectorStore.maxMarginalRelevanceSearch( "premium wireless over-ear headphones", { k: 4, fetchK: 12, lambda: 0.5 }, );</CODE></PRE><PRE><CODE>maxMarginalRelevanceSearch top 4: 1. AudioMax Pro 5 (over-ear-wireless, $349.99) 2. FlexFit Sport (neckband, $89.99) 3. AuroraSound Pro (over-ear-wireless, $379.99) 4. TuneCore 510 (on-ear-wireless, $49.99)</CODE></PRE><P>The top match is still the most relevant premium over-ear pair, but the next three deliberately span different form factors and price tiers. The shopper now sees an actual mix to compare, all from one method swap on the same vector store and the same product catalogue.</P><P>Use <CODE>similaritySearch</CODE> when the goal is the closest match. Reach for <CODE>maxMarginalRelevanceSearch</CODE> when the goal is <EM>coverage</EM>: recommendation strips, RAG context that spans multiple sub-topics, related-articles modules, or anywhere a single dominant theme would crowd out everything else.</P><H2 id="toc-hId-1031205690">Filtering that speaks JavaScript</H2><P>Semantic search is great. Semantic search <EM>plus</EM> structured filters is unbeatable. <CODE>HanaDB</CODE>’s filter syntax will feel instantly familiar to anyone who’s ever written a MongoDB query:</P><PRE><CODE>// In-stock electronics under $500 const filter = { $and: [ { in_stock: true }, { category: "electronics" }, { price: { $lt: 500 } }, ], }; const results = await store.similaritySearch("best audio quality", 10, filter);</CODE></PRE><P>And the result is exactly what the filter promised: only in-stock electronics, all under $500, ranked by audio quality:</P><PRE><CODE>[ { pageContent: 'Premium wireless headphones with spatial audio...', price: 349.99, in_stock: true }, { pageContent: 'Budget wireless earbuds with decent sound...', price: 49.99, in_stock: true }, { pageContent: 'Bluetooth over-ear headphones with active noise...', price: 249.99, in_stock: true }, { pageContent: 'Wireless noise-canceling headphones with 30-hour...',price: 299.99, in_stock: true } ]</CODE></PRE><P>You get the full operator suite (<CODE>$lt</CODE>, <CODE>$lte</CODE>, <CODE>$gte</CODE>, <CODE>$between</CODE>, <CODE>$in</CODE>, <CODE>$nin</CODE>, <CODE>$like</CODE>, <CODE>$contains</CODE>, <CODE>$and</CODE>, <CODE>$or</CODE>) all compiled down to native HANA SQL. No JavaScript-side filtering. No surprises at scale.</P><H2 id="toc-hId-834692185">Letting the LLM write its own filters</H2><P>Here’s where it gets fun. What if a user types <EM>“Show me AudioMax electronics under $400 with good ratings”</EM>? You <EM>could</EM> build a custom intent parser. Or you could let the LLM do it for free, with the <CODE>HanaTranslator</CODE>:</P><PRE><CODE>const retriever = SelfQueryRetriever.fromLLM({ llm, vectorStore: store, documentContents: "Product description for an e-commerce catalog", attributeInfo, // describe your metadata structuredQueryTranslator: new HanaTranslator(), }); await retriever.invoke("Show me AudioMax electronics under $400 with good ratings");</CODE></PRE><P>Behind the scenes, the LLM emits:</P><PRE><CODE>{ $and: [ { brand: "AudioMax" }, { category: "electronics" }, { price: { $lt: 400 } }, { rating: { $gte: 4 } } ] }</CODE></PRE><P>Natural language in. Structured filter out. Vector search on top. <STRONG>Zero glue code.</STRONG></P><H2 id="toc-hId-638178680">There’s more in the box</H2><P>What you’ve seen so far is the foundation. The library ships with three more capabilities, each big enough to deserve its own deep-dive:</P><UL><LI><STRONG><A href="https://github.com/SAP/langchainjs-integration-for-sap-hana-cloud/blob/main/examples/vectorstores/reranking.ts" target="_blank" rel="noopener nofollow noreferrer">Cross-encoding reranking</A></STRONG> with <CODE>HanaReranker</CODE>. Re-score your top candidates for surgical-precision results, the secret sauce behind great RAG and customer-facing search.</LI><LI><STRONG><A href="https://github.com/SAP/langchainjs-integration-for-sap-hana-cloud/blob/main/examples/graphs/basics.ts" target="_blank" rel="noopener nofollow noreferrer">Knowledge graphs &amp; SPARQL</A></STRONG> with <CODE>HanaRdfGraph</CODE>. Query RDF data, build natural-language Q&amp;A over relationships, and combine graph reasoning with vector search.</LI><LI><STRONG><A href="https://github.com/SAP/langchainjs-integration-for-sap-hana-cloud/blob/main/examples/vectorstores/internalEmbeddings.ts" target="_blank" rel="noopener nofollow noreferrer">Performance at scale</A></STRONG> with Map Merge, HNSW indexes, and specific metadata columns. Roughly 9× faster bulk inserts, sub-second search at million-document scale, and faster filtered queries.</LI></UL><P>Stay tuned for more updates and deep dives into each of these capabilities!</P><P><STRONG>Get the code:</STRONG> <A href="https://github.com/SAP/langchainjs-integration-for-sap-hana-cloud" target="_blank" rel="noopener nofollow noreferrer">SAP/langchainjs-integration-for-sap-hana-cloud</A> &nbsp;|&nbsp; <STRONG>Read the full guide:</STRONG> <A href="https://github.com/SAP/langchainjs-integration-for-sap-hana-cloud/blob/main/blogs/blog-post.md" target="_blank" rel="noopener nofollow noreferrer">blog-post.md</A></P> 2026-07-07T11:27:33.345000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/calculation-view-features-of-2026-qrc2/ba-p/14415862 Calculation View Features of 2026 QRC2 2026-07-08T14:50:53.797000+02:00 jan_zwickel https://community.sap.com/t5/user/viewprofilepage/user-id/239612 <P><SPAN>Within the time frame of 2026 Q2, several new calculation view features have been released in SAP Business Application Studio when connected to SAP HANA Cloud database QRC2. Some of these features are highlighted below. You can find examples that illustrate the individual features&nbsp;</SPAN><A href="https://github.com/SAP-samples/hana-cloud-learning/tree/main/CV_2026_QRC2_Selected_Calculation_View_Modeling_Features" target="_blank" rel="nofollow noopener noreferrer">here</A><SPAN>. An overview of features of other releases can be found&nbsp;</SPAN><A href="https://blogs.sap.com/2022/08/25/new-calculation-view-modeling-features-in-sap-hana-cloud/" target="_blank" rel="noopener noreferrer">here</A><SPAN>.</SPAN></P><H3 id="toc-hId-1946405822"><A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-performance-guide-for-developers/optimized-redeployment-of-calculation-views" target="_self" rel="noopener noreferrer">Optimized Redeployment</A></H3><P>When a calculation view is redeployed in a stacked scenario, all dependent objects in the same HDI container are normally dropped and sequentially recreated. This default behavior avoids cascading revalidations but comes with a cost: privileges on dependent objects are temporarily, or in cross-container scenarios potentially permanently, revoked, and the full recreate cycle adds significant deployment time.</P><P>Optimized redeployment addresses this. Instead of dropping all dependents upfront, a consistency check determines which dependent views are still valid after the change. Views unaffected by the redeployed calculation view are kept as-is and skipped entirely. This reduces both the time spent recreating unaffected views and the privilege reassignment overhead.</P><P>To enable optimized redeployment for calculation views, add the following parameter to the <CODE>hdi-deploy</CODE> invocation in <CODE>package.json</CODE>:</P><DIV class=""><PRE><CODE>"start": "node node_modules/@sap/hdi-deploy/deploy.js --parameter com.sap.hana.di.calculationview/optimized_redeploy=true"</CODE></PRE></DIV><P>A global parameter (<CODE>optimized_redeploy=true</CODE>) is also available and activates the behavior for all HDI object types that support it.<BR /><BR /></P><H3 id="toc-hId-1749892317"><A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-performance-guide-for-developers/optimized-redeployment-of-calculation-views" target="_self" rel="noopener noreferrer">Replace Instead of Drop and Recreate During Redeployment</A></H3><P>By default, calculation views are dropped and recreated during redeployment. Dropping a view revokes all privileges associated with it. Privileges defined in HDI roles of the same container are automatically regranted at the end of deployment, but direct privileges held by roles outside the current HDI container are permanently lost.</P><P>The new replace option substitutes the drop-and-recreate cycle with an in-place replacement. Because the object is never dropped, existing privileges are preserved throughout deployment. The benefit is most significant when combined with optimized redeployment: without it, dependent views are dropped before the root view is replaced, so they never benefit from the replace logic. With both options active, the replace strategy also applies to dependent views.</P><P>To enable the replace behavior, pass the following parameter:</P><DIV class=""><PRE><CODE>"start": "node node_modules/@sap/hdi-deploy/deploy.js --parameter optimized_replace=true"</CODE></PRE></DIV><P>Combining <CODE>optimized_replace=true</CODE> with <CODE>optimized_redeploy=true</CODE> gives the strongest deployment time and privilege preservation benefit.</P><P>&nbsp;</P><H3 id="toc-hId-1553378812"><A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-modeling-guide-for-sap-business-application-studio/supported-execution-hints" target="_self" rel="noopener noreferrer">Active/Active Read-Enabled Hints</A></H3><P>In an Active/Active (Read-Enabled) setup, query routing hints can now be configured at the individual calculation view level. This gives fine-grained control over whether a query is directed to the primary system or the read replica, allowing workload to be distributed and overall throughput to be increased.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="activeActiveHint.png" style="width: 764px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/420173i8E72DAB3B1B8B101/image-size/large?v=v2&amp;px=999" role="button" title="activeActiveHint.png" alt="activeActiveHint.png" /></span></P><P>&nbsp;</P><P>Hints are evaluated only on calculation views that are directly referenced in the query.&nbsp;</P><P>&nbsp;</P><H3 id="toc-hId-1356865307"><A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-modeling-guide-for-sap-business-application-studio/convert-join-nodes-to-non-equi-join-nodes" target="_self" rel="noopener noreferrer">Switching from a Join Node to a Non-Equi Join Node</A></H3><P>A Join node can now be converted directly to a Non-Equi Join node via the node's context menu, without needing to manually rebuild the join logic.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="convert.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/420174iDFE0CC3FB4F2FEA8/image-size/large?v=v2&amp;px=999" role="button" title="convert.png" alt="convert.png" /></span></P><P>During conversion you can choose how the resulting join condition is expressed: either graphically, based on the column display,</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="graphical.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/420175i61E31946FC8EC9A2/image-size/large?v=v2&amp;px=999" role="button" title="graphical.png" alt="graphical.png" /></span></P><P>or as an explicit join expression.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="expression.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/420176iE7BC947C16BB8D08/image-size/large?v=v2&amp;px=999" role="button" title="expression.png" alt="expression.png" /></span></P><P>This simplifies the refactoring of join logic and reduces manual rework. Note that features not supported by Non-Equi Join nodes, such as dynamic joins, text joins, and filter expressions, are removed automatically during the conversion.</P><H3 id="toc-hId-1160351802">Calculation View Status Dialog</H3><P>When opening a calculation view with certain properties, a status dialog now appears and reports details about the current view state. Where applicable, the dialog also offers direct actions to resolve issues, for example, prompting you to log in to Cloud Foundry when no active database connection is found.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="status.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/420178i974FE862FAC05D73/image-size/large?v=v2&amp;px=999" role="button" title="status.png" alt="status.png" /></span></P><P>The dialog covers four categories:</P><UL><LI><STRONG>Read-only mode:</STRONG>&nbsp;reasons why the calculation view was opened in read-only mode</LI><LI><STRONG>Database connection:</STRONG>&nbsp;current connection status to the database</LI><LI><STRONG><A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-modeling-guide-for-sap-business-application-studio/data-classification-for-calculation-views" target="_self" rel="noopener noreferrer">Data classification</A> inconsistencies:</STRONG>&nbsp;detected mismatches in the data classification setting</LI><LI><STRONG><A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-modeling-guide-for-sap-business-application-studio/calculation-views-for-sap-business-data-cloud" target="_self" rel="noopener noreferrer">BDC integration</A>:</STRONG>&nbsp;whether the calculation view is used with BDC integration mode</LI></UL><H3 id="toc-hId-963838297">&nbsp;</H3><H3 id="toc-hId-767324792"><A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-modeling-guide-for-sap-business-application-studio/define-mds-cube-based-on-elements-in-calculation-view" target="_self" rel="noopener noreferrer">Calculated Measures in MDS Cubes</A></H3><P>Calculated measures can now be added to MDS Cubes and are evaluated at query runtime. This makes it possible to execute calculations at query runtime.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="calculatedMeasures.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/420183iC91644058E048978/image-size/large?v=v2&amp;px=999" role="button" title="calculatedMeasures.png" alt="calculatedMeasures.png" /></span></P><P>If a calculated measure cannot be translated into a valid MDS formula, deployment fails with a corresponding error message.</P> 2026-07-08T14:50:53.797000+02:00 https://community.sap.com/t5/technology-blog-posts-by-members/media-handling-in-sap-cap-mixing-static-and-uploaded-images-in-fiori/ba-p/14432155 Media Handling in SAP CAP: Mixing Static and Uploaded Images in Fiori Elements V4 2026-07-09T20:15:58.012000+02:00 Shubham_kumar_sap https://community.sap.com/t5/user/viewprofilepage/user-id/2084304 <P>Hi everyone,</P><H3 id="toc-hId-1948156736">Introduction</H3><P>When building applications with the SAP Cloud Application Programming (CAP) model and SAP Fiori Elements V4, handling media files is a common requirement. Usually, it is straightforward: first define a <FONT color="#00FF00">LargeBinary&nbsp;</FONT>field, and the framework generates an upload control.</P><P>However, things get complicated when we need a single table column to dynamically display images from <I>two different sources</I>:</P><OL><LI><P>Static default images (loaded via CSV during deployment).</P></LI><LI><P>Custom images uploaded by users via the UI.</P></LI></OL><P>This guide covers the complete end-to-end setup to achieve this:</P><UL><LI><P><STRONG>Data Modeling</STRONG> — Defining the schema with a "Traffic Cop" virtual field.</P></LI><LI><P><STRONG>UI Annotations</STRONG> — Exposing the right fields to the List Report and Object Page.</P></LI><LI><P><STRONG>Backend Logic</STRONG> — Intercepting OData V4 requests in Node.js to bypass UI optimizations.</P></LI><LI><P><STRONG>HANA Deployment</STRONG> — Resolving common driver version conflicts during deployment.</P></LI></UL><H3 id="toc-hId-1751643231">Why I Wrote This</H3><P>Recently, I was developing an application on the SAP Business Technology Platform where users needed to browse a catalog of items. Some items had static placeholder images mapped to local CSV paths, while others had high-res photos explicitly uploaded by users.</P><P>When I attempted to route these dynamically using a virtual field, my Fiori V4 List Report stubbornly displayed blank "No Image" placeholders. After hours of debugging the OData V4 network payload, I realized Fiori was optimizing my request by stripping out the hidden columns my backend logic relied on. I couldn't find a definitive guide covering this specific Fiori V4 optimization trap, so I'm sharing the bulletproof workaround I built.</P><P>&nbsp;</P><H3 id="toc-hId-1555129726">First, Let's Understand the Fiori V4 Optimization Trap</H3><P>Before looking at the code, it helps to understand why the standard approach fails in Fiori Elements V4.</P><P>Fiori V4 is aggressively optimized for bandwidth. When it loads a List Report, the OData <FONT color="#00FF00">$select</FONT>&nbsp;query <I>only</I> asks for the columns explicitly visible on the screen. If the backend relies on a virtual field (e.g., <FONT color="#00FF00">displayImageUrl</FONT>) that calculates its value based on hidden fields (like <FONT color="#00FF00">mediaType&nbsp;</FONT>or <FONT color="#00FF00">ImageUrl</FONT>), those hidden fields arrive at the backend handler as <FONT color="#00FF00">undefined</FONT>. We cannot calculate a dynamic URL if the framework refuses to fetch the underlying data.</P><P>To fix this, we don't fight the UI. Instead, we let Fiori request whatever it wants, and we use a custom CAP backend hook to silently query the database for the missing fields before delivering the final payload to the browser.</P><P>&nbsp;</P><H3 id="toc-hId-1358616221">Architecture Overview</H3><P>This solution relies on a clean separation between the database schema, the UI annotations, and the backend Node.js service.</P><pre class="lia-code-sample language-javascript"><code>project-root/ │ ├── db/ │ ├── data/ │ │ └── mediahandling.db-Image.csv ← Initial data linking to static image paths │ └── schema.cds ← Defines the binary, static, and virtual fields │ ├── app/ │ ├── mediahandling/ ← Your Fiori Elements application │ │ └── webapp/ │ │ └── images/ │ │ ├── image1.jpg ← The actual static placeholder files │ │ └── image2.jpg │ └── annotations.cds ← Maps the UI controls for Fiori Elements │ └── srv/ ├── catalogService.cds ← Exposes the OData service └── catalogService.js ← Contains the custom 'READ' interceptor logic</code></pre><P>&nbsp;</P><H3 id="toc-hId-1162102716">Step 1: The Data Model (schema.cds)</H3><P>We need an entity that stores both the binary file and the static URL, alongside a virtual field to decide which one to render.</P><P><I>Critical Detail:</I> Ensure the <FONT color="#00FF00"><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/1445379">@ui</a>.IsImageURL: true</FONT> annotation is applied directly to the virtual field here in the schema. This instructs Fiori V4 to natively allocate an avatar control in the List Report.</P><pre class="lia-code-sample language-javascript"><code>namespace mediahandling.db; using { cuid, managed } from '@sap/cds/common'; entity Image: cuid, managed { Name : String(100); Description : String(255); // Holds the static CSV path (e.g., 'images/default.jpg') ImageUrl : String; // The actual uploaded file data @Core.MediaType : mediaType image : LargeBinary; // The file extension/type (e.g., 'image/jpeg') @Core.IsMediaType: true mediaType : String; // The "Traffic Cop" virtual field for Fiori UI display @ui.IsImageURL: true virtual displayImageUrl : String; }</code></pre><P>&nbsp;</P><H3 id="toc-hId-965589211">Step 2: Initializing the Mock Data (CSV)</H3><P>To test this, we need to load initial records into the database. Create a file named <FONT color="#00FF00">mediahandling.db-Image.csv&nbsp;</FONT>inside the&nbsp;<FONT color="#00FF00">db/data/</FONT> folder.</P><PRE><CODE>ID;Name;Description;ImageUrl 22222222-2222-2222-2222-222222222201;Sunset Boulevard;A beautiful vibrant sunset over the city boulevard.;images/image1.jpg 22222222-2222-2222-2222-222222222202;Mountain Peak;The highest snow-capped peak in the northern mountain range.;images/image2.jpg</CODE></PRE><P>&nbsp;</P><H3 id="toc-hId-769075706">Step 3: UI Annotations (annotations.cds)</H3><P>In the UI annotations, we expose the virtual <FONT color="#00FF00">displayImageUrl</FONT> field&nbsp;in the <FONT color="#00FF00">UI.LineItem</FONT>&nbsp;(List Report) so the dynamic image renders in the table. We expose the actual <FONT color="#00FF00">image&nbsp;</FONT>(<FONT color="#00FF00">LargeBinary</FONT>) field inside the <FONT color="#00FF00">UI.FieldGroup</FONT>&nbsp;(Object Page) so the user gets a file upload prompt when creating or editing a record.</P><pre class="lia-code-sample language-javascript"><code>using MyService as service from '../../srv/catalogService'; annotate service.Image with @( UI.FieldGroup #GeneratedGroup : { $Type : 'UI.FieldGroupType', Data : [ { $Type : 'UI.DataField', Label : 'Name', Value : Name }, { $Type : 'UI.DataField', Label : 'Description', Value : Description }, { $Type : 'UI.DataField', Label : 'Upload Image', Value : image } ], }, UI.Facets : [ { $Type : 'UI.ReferenceFacet', ID : 'GeneratedFacet1', Label : 'General Information', Target : '@UI.FieldGroup#GeneratedGroup', }, ], UI.LineItem : [ { $Type : 'UI.DataField', Label : 'Name', Value : Name }, { $Type : 'UI.DataField', Label : 'Image', Value : displayImageUrl }, { $Type : 'UI.DataField', Label : 'Description', Value : Description } ], UI.HeaderInfo:{ $Type : 'UI.HeaderInfoType', TypeName : 'Image Record', TypeNamePlural : 'Image Records', Title : { $Type: 'UI.DataField', Value: ID }, Description : { $Type: 'UI.DataField', Value: Name }, ImageUrl : displayImageUrl } );</code></pre><P>&nbsp;</P><H3 id="toc-hId-572562201">Step 4: The Backend Interceptor (catalogService.js)</H3><P>This is where the magic happens. We intercept the <FONT color="#00FF00">after('READ')</FONT> event. If we detect that Fiori stripped our necessary columns (<FONT color="#00FF00">mediaType</FONT> and <FONT color="#00FF00">ImageUrl</FONT>), we immediately execute a lightweight database query to fetch them for the visible rows. Finally, we route the URL accordingly.</P><pre class="lia-code-sample language-javascript"><code>const cds = require('@sap/cds'); module.exports = cds.service.impl(async function() { // Dynamically grab the actual service base path (e.g., /odata/v4/my) const srvPath = this.path; this.after('READ', 'Image', async (data, req) =&gt; { const images = Array.isArray(data) ? data : [data]; if (images.length === 0) return; // If Fiori stripped the columns to save bandwidth, we step in. if (images[0].mediaType === undefined &amp;&amp; images[0].ImageUrl === undefined) { // Grab the IDs of the rows currently visible on the screen const ids = images.map(img =&gt; img.ID); // Query the database directly to fetch the missing fields const dbData = await cds.tx(req).run( SELECT.from(req.target) .columns('ID', 'mediaType', 'ImageUrl', 'IsActiveEntity') .where({ ID: { 'in': ids } }) ); // Stitch the missing data back into the original payload images.forEach(img =&gt; { const match = dbData.find(d =&gt; d.ID === img.ID); if (match) { img.mediaType = match.mediaType; img.ImageUrl = match.ImageUrl; img.IsActiveEntity = match.IsActiveEntity; } }); } // Execute the Traffic Cop logic (guaranteed to have the data) images.forEach(image =&gt; { if (image.mediaType) { // Route to the OData media stream for uploaded binaries const isActive = image.IsActiveEntity !== false; image.displayImageUrl = `${srvPath}/Image(ID=${image.ID},IsActiveEntity=${isActive})/image`; } else if (image.ImageUrl) { // Route to the static CSV path image.displayImageUrl = image.ImageUrl; } else { // Fallback placeholder image.displayImageUrl = 'https://dummyimage.com/200x200/cccccc/000000.png&amp;text=No+Image'; } }); }); });</code></pre><H3 id="toc-hId-376048696">&nbsp;</H3><H3 id="toc-hId-179535191">Output:</H3><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Screenshot 2026-07-02 175031.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428527iC2F9FBC6B23C85D4/image-size/large?v=v2&amp;px=999" role="button" title="Screenshot 2026-07-02 175031.png" alt="Screenshot 2026-07-02 175031.png" /></span></P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Screenshot 2026-07-02 180108.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428530iFBA9B183BEE9527C/image-size/large?v=v2&amp;px=999" role="button" title="Screenshot 2026-07-02 180108.png" alt="Screenshot 2026-07-02 180108.png" /></span></P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Screenshot 2026-07-02 180226.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428531i1438E6232EF8D48A/image-size/large?v=v2&amp;px=999" role="button" title="Screenshot 2026-07-02 180226.png" alt="Screenshot 2026-07-02 180226.png" /></span></P><P><span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Screenshot 2026-07-02 180242.png" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/428532i2F7FFCCD58978D4C/image-size/large?v=v2&amp;px=999" role="button" title="Screenshot 2026-07-02 180242.png" alt="Screenshot 2026-07-02 180242.png" /></span></P><P>&nbsp;</P><H3 id="toc-hId--92209683">Note: Deploying to SAP HANA Cloud</H3><P>When transitioning from a local SQLite setup to SAP HANA Cloud, we may encounter an NPM dependency error when starting the hybrid profile.</P><P>If the project is running <FONT color="#00FF00"><a href="https://community.sap.com/t5/user/viewprofilepage/user-id/2302137">@SAP</a>/cds</FONT> version 9, but the latest <FONT color="#00FF00">@cap-js/hana</FONT> driver expects version 10, the standard <FONT color="#00FF00">npm install</FONT> will fail with an <FONT color="#00FF00">ERESOLVE</FONT>&nbsp;conflict. We can safely bypass this strict peer dependency check and link the driver by running the following:</P><pre class="lia-code-sample language-bash"><code>npm install @cap-js/hana --legacy-peer-deps</code></pre><DIV class=""><DIV class=""><DIV class=""><DIV class="">Once installed,&nbsp;<FONT color="#00FF00">cds watch --profile hybrid&nbsp;</FONT>command will connect to the HANA instance seamlessly.</DIV><DIV class="">&nbsp;</DIV><DIV class=""><H3 id="toc-hId--288723188">Conclusion</H3><P>By shifting the data-retrieval logic into the Node.js backend rather than relying on messy UI workarounds or forced hidden columns, we maintain the strict performance benefits of Fiori Elements V4 while delivering a seamless, dynamic user experience.</P></DIV></DIV></DIV></DIV> 2026-07-09T20:15:58.012000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/sap-hana-cloud-release-june-2026-round-up/ba-p/14440351 SAP HANA Cloud release June 2026 Round-up 2026-07-15T16:05:19.094000+02:00 andreamiranda https://community.sap.com/t5/user/viewprofilepage/user-id/135788 <P><SPAN>Dear SAP HANA Cloud Enthusiasts,</SPAN><BR /><BR />We are excited to share a collection of the latest videos, blog posts, and resources highlighting the SAP HANA Cloud Q2 2026 release.</P><P>&nbsp;</P><TABLE border="1" width="100%"><TBODY><TR><TD width="50.129533678756474%" height="222px"><A href="https://www.youtube.com/watch?v=FsTR3NR1he0&amp;list=PL3ZRUb1AKkpTDZQgENtRcupp6vsNg8NHN&amp;index=2" target="_blank" rel="noopener nofollow noreferrer"><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="playTeaserpng.png" style="width: 400px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/432945iADFAAE345978F76B/image-size/medium?v=v2&amp;px=400" role="button" title="playTeaserpng.png" alt="playTeaserpng.png" /></span></A></TD><TD width="49.870466321243526%" height="222px"><H4 id="toc-hId-2078105312"><STRONG>What’s New Teaser</STRONG></H4>Explore the latest innovations in SAP HANA Cloud with Lead Product Manager Thomas Hammer, as he shares his top highlights from the newest release in this engaging teaser.<BR /><BR /><A href="https://www.youtube.com/watch?v=FsTR3NR1he0&amp;list=PL3ZRUb1AKkpTDZQgENtRcupp6vsNg8NHN&amp;index=2" target="_blank" rel="nofollow noopener noreferrer">Watch it now on YouTube.</A></TD></TR><TR><TD width="50.129533678756474%" height="250px"><H4 id="toc-hId-1881591807"><STRONG>What’s New blogpost</STRONG></H4>Intrigued by the teaser? Explore our "What’s New in SAP HANA Cloud in June 2026" blogpost for an in-depth look at the innovations and find valuable links to further demos and content.<BR /><BR /><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/what-s-new-in-sap-hana-cloud-july-2026/ba-p/14419199" target="_blank">Read it here!</A></TD><TD width="49.870466321243526%" height="250px"><A href="https://community.sap.com/t5/technology-blog-posts-by-sap/what-s-new-in-sap-hana-cloud-july-2026/ba-p/14419199" target="_blank"><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="blogpostSS.png" style="width: 400px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/432946iB1F55C6248D7F214/image-size/medium?v=v2&amp;px=400" role="button" title="blogpostSS.png" alt="blogpostSS.png" /></span></A></TD></TR><TR><TD width="50.129533678756474%" height="250px"><A href="https://www.youtube.com/watch?v=QrGR38jGGZo&amp;list=PL3ZRUb1AKkpTDZQgENtRcupp6vsNg8NHN&amp;index=1" target="_blank" rel="noopener nofollow noreferrer"><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="PlayWebinar.png" style="width: 400px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/432947i894D8E6CA32AB861/image-size/medium?v=v2&amp;px=400" role="button" title="PlayWebinar.png" alt="PlayWebinar.png" /></span></A></TD><TD width="49.870466321243526%" height="250px"><H4 id="toc-hId-1685078302">What's New Webinar</H4>Prefer watching over reading? Our 'What’s New' webinar is just for you! Join our Product experts for an in-depth view of the latest features. Available to watch anytime&nbsp;<A href="https://www.youtube.com/watch?v=QrGR38jGGZo&amp;list=PL3ZRUb1AKkpTDZQgENtRcupp6vsNg8NHN&amp;index=1" target="_blank" rel="noopener nofollow noreferrer">here</A>!</TD></TR></TBODY></TABLE><H3 id="toc-hId-1359482078">&nbsp;</H3><DIV class=""><HR /><SPAN>Don’t miss out on all the content and remember to&nbsp;</SPAN><A href="https://community.sap.com/topics/hana" target="_blank">follow us in the SAP HANA Community.</A></DIV><P>Remember to check our content following the # whatsnewinsaphanacloud tag:<SPAN>&nbsp;</SPAN><A href="https://community.sap.com/t5/tag/whatsnewinsaphanacloud/tg-p/board-id/technology-blog-sap" target="_blank">here</A><BR /><BR /><SPAN>Don’t forget to subscribe and follow SAP HANA Cloud on&nbsp;</SPAN><A href="https://www.youtube.com/playlist?list=PL3ZRUb1AKkpTDZQgENtRcupp6vsNg8NHN" target="_blank" rel="nofollow noopener noreferrer">YouTube</A><SPAN>&nbsp;to always stay up-to-date regarding the most recent innovations in SAP HANA Cloud.</SPAN><BR /><SPAN>&nbsp;</SPAN><BR /><SPAN>All the best,</SPAN><BR /><BR /><STRONG>Andrea on behalf of the SAP HANA Cloud team</STRONG></P><P><STRONG><A class="" href="https://community.sap.com/t5/c-khhcw49343/SAP+HANA+Cloud%25252C+SAP+HANA+database/pd-p/ada66f4e-5d7f-4e6d-a599-6b9a78023d84" target="_blank">#SAP HANA Cloud, SAP HANA database</A><SPAN>&nbsp;</SPAN>&nbsp;<SPAN>&nbsp;#</SPAN><A class="" href="https://community.sap.com/t5/c-khhcw49343/SAP+HANA+Cloud/pd-p/73554900100800002881" target="_blank">SAP HANA Cloud</A><SPAN>&nbsp;</SPAN>&nbsp;</STRONG></P><P><STRONG>#whatsnewinsaphanacloud</STRONG></P> 2026-07-15T16:05:19.094000+02:00 https://community.sap.com/t5/technology-blog-posts-by-sap/from-preview-to-production-get-started-with-the-sap-hana-cloud-s-agentic/ba-p/14440515 From Preview to Production: Get Started with the SAP HANA Cloud's Agentic Data Exploration Tools 2026-07-15T21:25:57.846000+02:00 shabana https://community.sap.com/t5/user/viewprofilepage/user-id/38259 <P>At SAP TechEd 2025, we shared a vision: what if exploring data felt less like archaeology and more like a conversation? We introduced two capabilities designed to make that real and the response told us we were onto something. Fast forward to today. That vision is now in production. <STRONG>SAP HANA Cloud's data exploration capabilities are generally available, and customers can start using them right now.</STRONG></P><P>If you followed the TechEd announcement, you will remember these capabilities as the <A href="https://community.sap.com/t5/technology-blog-posts-by-sap/sap-hana-cloud-becomes-agentic-introducing-discovery-agent-amp-data-agent/ba-p/14257394" target="_blank">Discovery Agent and the Data Agent</A>. As they matured from prototype to production, so did our thinking about what to call them.</P><P>They are now formally known as the <STRONG>Database Object Discovery Tool</STRONG> and the <STRONG>Data Retrieval Tool</STRONG> - together, the <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-administration-guide/data-exploration-with-generative-ai?locale=en-US&amp;version=LATEST" target="_blank" rel="noopener noreferrer"><STRONG>SAP HANA Cloud data exploration tools</STRONG></A>. The rename reflects a natural evolution: as these capabilities became production-grade, composable building blocks that any agent, application, or workflow can invoke, the "tool" framing became the more precise one. The underlying power is the same.&nbsp;</P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="Note: **Upcoming Agents. Subject to change as we evolve our roadmap" style="width: 532px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/433037iB2C93186F26236A1/image-size/large?v=v2&amp;px=999" role="button" title="Blog_D&amp;DTools_Pic.png" alt="Note: **Upcoming Agents. Subject to change as we evolve our roadmap" /><span class="lia-inline-image-caption" onclick="event.preventDefault();">Note: **Upcoming Agents. Subject to change as we evolve our roadmap</span></span></P><P><STRONG>Why This Matters - For Everyone</STRONG></P><P>Whether you are a developer building an AI-powered application, or a business stakeholder who simply wants answers from data, these tools close a gap that has existed for a long time.</P><P>For business users and decision makers:</P><UL><LI><STRONG>You no longer need to know where data lives to find it. </STRONG>The Database Object Discovery Tool searches across tables, views, and relationships based on what you describe, or query about.</LI><LI><STRONG>You no longer need SQL to get results. </STRONG>The Data Retrieval Tool translates a plain language request into a query and returns the answer, no manual lookup, no schema expertise required.</LI><LI><STRONG>The experience feels conversational, </STRONG>even though sophisticated discovery and query generation is happening beneath the surface.</LI></UL><P>For developers and architects:</P><UL><LI><STRONG>Production-grade interfaces: </STRONG>stored-procedure-based (<A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-administration-guide/ai-object-retrieval-procedure?locale=en-US&amp;version=LATEST" target="_blank" rel="noopener noreferrer">AI_OBJECT_DISCOVERY</A> and <A href="https://help.sap.com/docs/hana-cloud-database/sap-hana-cloud-sap-hana-database-administration-guide/ai-data-retrieval-procedure?locale=en-US&amp;version=LATEST" target="_blank" rel="noopener noreferrer">AI_DATA_RETRIEVAL</A>) that integrate cleanly into any application or agent workflow.</LI><LI><STRONG>Composable by design: </STRONG>wire them into an MCP server or call them directly from your own tooling &amp; agents.</LI><LI><STRONG>Context inherited automatically: </STRONG>the Custom Database Objects Knowledge Graph continues to serve as the semantic backbone, so your applications gain that contextual understanding without you having to build it.</LI></UL><P><STRONG>Getting Started </STRONG></P><P>There is no additional infrastructure to provision. If you have access to SAP HANA Cloud with the Natural Language Processing &amp; Triple Store services enabled, and an SAP AI Core subscription, you have access to these tools today.</P><P>The pattern is straightforward:</P><OL><LI><STRONG>Discover -&nbsp;</STRONG>Describe what you are looking for in natural language. The Database Object Discovery Tool searches the HANA catalog's knowledge graph &amp; RAG layer and returns the relevant tables, views, and columns.</LI><LI><STRONG>Retrieve -&nbsp;</STRONG>Pass that context to the Data Retrieval Tool along with your question. It generates and executes the appropriate query and returns results.</LI></OL><P><STRONG>Try It Yourself -&nbsp;&nbsp;<A href="https://developers.sap.com/tutorials/hana-cloud-data-exploration-tools.html" target="_blank" rel="noopener noreferrer"><U>Explore Data with SAP HANA Cloud Data Exploration Tools</U></A></STRONG></P><P>The SAP Developer Tutorials has a step by step guide that walks you through setting up both tools and running your first natural-language query against SAP HANA Cloud. It covers the key features that make HANA Cloud well-suited for business AI, and gets you hands-on quickly.</P><P><span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="SAP HANA Cloud Central Tooling" style="width: 999px;"><img src="https://community.sap.com/t5/image/serverpage/image-id/433602i8299827FB777A805/image-size/large?v=v2&amp;px=999" role="button" title="Get Started pic.png" alt="SAP HANA Cloud Central Tooling" /><span class="lia-inline-image-caption" onclick="event.preventDefault();">SAP HANA Cloud Central Tooling</span></span></P><P>&nbsp;</P><DIV class=""><DIV class=""><STRONG>What's Next: Taking the Tools Further with MCP</STRONG></DIV><DIV class="">&nbsp;</DIV><DIV class="">These tools are built to be composable and one of the most powerful ways to use them is through the Model Context Protocol (MCP). In the next post in this series, we will walk through how to spin up a local MCP client that connects directly to the Database Object Discovery Tool and the Data Retrieval Tool, and what that unlocks for agentic AI workflows on top of SAP HANA Cloud.</DIV><DIV class="">If you have been curious about how SAP HANA Cloud fits into the broader agentic architecture picture at the protocol level that one is for you.</DIV><DIV class="">The Data Exploration Tools are more than a new capability, they mark SAP HANA Cloud's evolution toward an agentic database that not only manages data, but also enables users and AI agents to understand, explore, and leverage it intelligently. This is the direction SAP HANA Cloud is moving and there is more to come.</DIV></DIV> 2026-07-15T21:25:57.846000+02:00