SShortSingh.
Back to feed

Databricks FILE Type Works on FSx for ONTAP S3 Access Points, But Reads Fail

0
·3 views

A developer tested Databricks' beta FILE column type against files stored on Amazon FSx for NetApp ONTAP, accessed via S3 Access Points, publishing findings on August 12, 2026. While registering files against an FSx for ONTAP S3 Access Point succeeded, read operations consistently failed due to restrictions in the session policy attached to temporary credentials issued by Unity Catalog. The FILE type itself functioned correctly, but users must choose between FILE EXTERNAL and FILE MANAGED before ingestion, as the decision cannot be reversed afterward. Object tag values are effectively limited to ASCII, with most CJK character strings being rejected — a behavior AWS Support has escalated to its service team as a potential defect. Additionally, the operator support is inconsistent, with GROUP BY and DISTINCT accepted but equality checks and ORDER BY rejected for FILE type columns.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingHacker News ·

UK Met Office Launches Glacier Tracking Tool on Climate Dashboard

The UK Met Office has added a glacier monitoring section to its online Climate Dashboard. The tool provides data visualizations tracking the state and changes of glaciers globally. Glaciers are key indicators of climate change, making their inclusion on such platforms scientifically significant. The dashboard is publicly accessible and aims to present climate data in a transparent, understandable format.

0
ProgrammingDEV Community ·

How to Self-Host Langfuse LLM Observability Platform Using Docker Compose

Langfuse is an open-source observability platform designed to monitor LLM applications by tracking traces, token usage, costs, and providing debugging analytics for AI workflows. A technical guide published on DEV Community outlines how to deploy Langfuse on a Linux server using Docker Compose, combining PostgreSQL, ClickHouse, Redis, and S3-compatible object storage. The setup is secured with Traefik as a reverse proxy and uses Let's Encrypt for automated TLS certificate management. Deployment requires a minimum of 4 vCPUs and 16GB RAM, a configured domain A record, and six randomly generated secrets for securing database and application credentials. Once running, the platform allows developers to send real traces through the stack and monitor production AI application behaviour from a self-hosted environment.

0
ProgrammingDEV Community ·

Controlled Vocabularies, Taxonomies, and Ontologies: Know What You Actually Need

In knowledge management, controlled vocabularies, taxonomies, and ontologies represent three distinct levels of data structuring, each roughly an order of magnitude more complex than the previous. A controlled vocabulary is simply a fixed list of agreed-upon terms with definitions, identifiers, and statuses — enough to enable consistent filtering and reporting. Adding a broader/narrower hierarchy to those terms creates a taxonomy, which enables roll-up queries and faceted navigation but introduces challenges like hierarchy disputes and non-tree-shaped domains. Ontologies go further by defining classes, properties, and logical axioms that allow machines to infer new facts, but require specialized modeling expertise and reasoning infrastructure. The article cautions that the term 'ontology' is frequently misused to describe all three levels, and most teams seeking one actually need nothing more than a well-maintained list of forty agreed-upon terms.

0
ProgrammingDEV Community ·

How to Safely Export and Verify ML Models Using ONNX Runtime

ONNX Runtime allows a single model artifact to run across CPU, GPU, and various accelerators via one API, eliminating much of the per-platform engineering work. However, the export process can silently introduce errors, including frozen control flow, operator decomposition into approximations, and numeric drift. Developers must carefully configure dynamic axes during export to avoid inference failures on variable-sized inputs, and should validate the exported graph's node count as an early warning of inefficient decomposition. Numerical accuracy should be verified by running both the original and exported models on identical inputs and comparing outputs within an explicit tolerance threshold. Opset versioning and provider-level operator support are additional compatibility concerns that can cause parts of a model to run on unintended hardware without any error being raised.