New pysmartdatamodels Version 0.8.1.0: Refreshed Data (reduced size 20 times with same functionality)

We are happy to announce a new release of our Python package, pysmartdatamodels, version 0.8.1.0, with a refreshed attribute database, a smaller download, and a cleanup of some leftover dead code.

What’s New in This Version?

Refreshed Data, End to End

The package’s bundled data — the official list of data models, the attribute metadata, and the full attribute database — had gone stale. Fixing this surfaced two real issues in our own publication pipeline, now resolved:

  • One data model repository had accidentally been left private, which was silently breaking the daily metadata regeneration every time it ran.
  • The database backing the attribute export (see below) was far less complete than it should have been. It’s now fully backfilled: 164,926 attribute records across all 82 subjects.

As always, you don’t need to upgrade the package just to get fresh data — calling update_data() pulls the latest information directly:

from pysmartdatamodels import pysmartdatamodels as sdm
sdm.update_data()

Smaller Install, Compressed Attribute Database

The full attribute database is now shipped gzip-compressed inside the package (~5MB instead of ~110MB) — the installed package footprint drops from over 100MB to under 10MB. This is fully transparent if you use the documented functions (load_all_attributes(), description_attribute(), etc.) — nothing changes in how you call them.

The same compressed file is now also available for direct download, alongside the existing uncompressed one, from the attribute database export page.

Cleanup

Removed a small piece of dead code in generate_sql_schema() left over from an earlier implementation — no change in behavior, just a cleaner function. Thanks to the community member who originally reported it.

Get the Latest Version

Update your installation with:

pip install --upgrade pysmartdatamodels

➡️ Get the package on PyPI

New pysmartdatamodels Version 0.8.0.11: More Reliable SQL Schema Generation

We are happy to announce a new release of our Python package, pysmartdatamodels, version 0.8.0.11, focused on making the SQL schema generation function more robust, fixing how the package exposes its functions, and keeping our GitHub source and published package in sync.

What’s New in This Version?

More Reliable SQL Schema Generation

The generate_sql_schema() function, which turns a data model’s model.yaml into a ready-to-use PostgreSQL CREATE TABLE statement, received three fixes:

  • Column names are now properly quoted, so attribute names that are reserved words or mixed case no longer produce invalid SQL.
  • Properties using anyOf are now supported, in addition to oneOf.
  • The function no longer emits a duplicate id column in the generated schema.

Example:

from pysmartdatamodels import pysmartdatamodels as sdm

sql = sdm.generate_sql_schema("https://raw.githubusercontent.com/smart-data-models/dataModel.Weather/master/WeatherForecast/model.yaml")
print(sql)

More Usable Imports

You can now import the package’s main functions directly from the top level, instead of always going through the submodule:

from pysmartdatamodels import generate_sql_schema, load_all_datamodels

The documented usage pattern below continues to work exactly as before:

from pysmartdatamodels import pysmartdatamodels as sdm

Repository Housekeeping

The source code published on GitHub and the package published on PyPI had drifted apart over the last few releases. With this version, both are back in sync, so what you see in the repository is exactly what you get when you pip install the package.

Get the Latest Version

Update your installation with:

pip install --upgrade pysmartdatamodels

➡️ Get the package on PyPI

Our Commitment to Open and Interoperable Data

We are committed to making data interoperability easier for everyone. Thanks to everyone in the community who reports issues and helps us keep improving pysmartdatamodels!

New candidate models: SIMPLE’s Business Service extensions to FEDeRATED

Nine new candidate data models are now available in the Smart Data Models Candidates repository, covering business transactions and logistics services for road and rail freight, translated from the SIMPLE project’s extensions to the FEDeRATED Business Service ontology.

The models cover the full lifecycle of a freight movement: Consignment and Shipment (with delivery time windows and an eFTI UID for EU-Gate interoperability), TransportService and CargoHandlingService, and five order types used at cargo terminals — CargoHandlingOrder, DischargeOrder, GateInOrder, GateOutOrder and ShuntingOrder.

These models are grounded in the ontology published at vocab.plataformasimple.es and its imported FEDeRATED Business Service base ontology, with each entity’s schema reflecting the full set of inherited attributes — including required fields drawn directly from the base ontology’s own cardinality constraints.

Explore the models: standards/simple-federated-businessservice in the Candidates repository.

New candidate model: SIMPLE’s Digital Twin extension to FEDeRATED (CMR)

A new candidate data model is now available in the Smart Data Models Candidates repository: CMR, the international road freight consignment note (“carta de porte”), translated from the SIMPLE project’s extension to the FEDeRATED Digital Twin ontology.

CMR documents are the standard legal paperwork accompanying road freight shipments across most of Europe. The model captures its identifying attributes (document ID, version, type, MRN number) and external identifiers, ready to be linked from a business transaction.

Explore the model: standards/simple-federated-digitaltwin in the Candidates repository.

New candidate models: SIMPLE’s Event extensions to FEDeRATED

Twenty new candidate data models are now available in the Smart Data Models Candidates repository, covering road and rail logistics events translated from the SIMPLE project’s extensions to the FEDeRATED Event ontology.

The models cover terminal movements (GateInEvent, GateOutEvent), generic transport (VehicleArrivalEvent, VehicleDepartureEvent, PassingEvent), workforce time tracking (ClockInConfirmationEvent, ClockOutConfirmationEvent), incident reporting (IssueEvent), and the complete eCMR (electronic Consignment Note) document lifecycle — creation, sender/carrier/consignee signatures and comment edits, destination and load changes, cancellation, and signature invalidation.

Every event links back to the business transaction, digital twin and location it concerns, so these events compose directly with the recently added simple-federated-businessservice candidate models.

Explore the models: standards/simple-federated-event in the Candidates repository.

New candidate models: SIMPLE’s Logistic Roles extensions to FEDeRATED

Two new candidate data models are now available in the Smart Data Models Candidates repository, covering role-based context for logistics locations and drivers, translated from the SIMPLE project’s extensions to the FEDeRATED Logistic Roles ontology.

Driver represents a person driving a vehicle, referenced directly by the sibling Business Service candidate’s TransportService entity. LocationRole captures the role a physical location plays in a specific transport context — port of loading, place of delivery, gate-in location, and 15 other role types — tied to the transport means and event it relates to.

These models are grounded in the ontology published at vocab.plataformasimple.es, which itself republishes the FEDeRATED project’s Logistic Roles taxonomy in full alongside SIMPLE’s own additions.

Explore the models: standards/simple-federated-logisticroles in the Candidates repository.

New candidate model: SIMPLE’s Physical Infrastructure extension to FEDeRATED

A new Location candidate model is now available in the Smart Data Models Candidates repository, translated from the SIMPLE project’s extension to the FEDeRATED Physical Infrastructure ontology.

The model adds the contact details (address, email, telephone) SIMPLE requires on any physical logistics location — terminals, warehouses, ports and similar facilities — alongside location-role and containment relationships carried over from the base ontology.

Explore the model: standards/simple-federated-physicalinfrastructure in the Candidates repository.

New candidate model: SIMPLE’s Classifications extensions to FEDeRATED

A new candidate data model, Classifications, is now available in the Smart Data Models Candidates repository, translated from the SIMPLE project’s extensions to the FEDeRATED Classifications ontology.

Classifications is the controlled-vocabulary backbone shared across FEDeRATED’s logistics ontologies: document types, cargo/container/package/equipment types, dangerous goods and nature-of-cargo codes, transport modality and movement types, seal condition codes, and TEN-T core network corridor names. Rather than one entity type per code list, this candidate models a single generic Classifications entity — each instance a code value tagged with the specific scheme it belongs to — keeping the model proportionate to what is genuinely a taxonomy, not a set of independent entity types.

SIMPLE’s own contribution: CMR, the international road consignment note, added as a new DocumentType code value.

Explore the model: standards/simple-federated-classifications in the Candidates repository.

European Union Agency for Railways (ERA) available in the candidates repository

Smart Data Models’ Candidates repository now includes four new candidate standards translated from the European Union Agency for Railways (ERA) Ontology v3.3.3 — the shared data model behind the EU’s four railway interoperability registers. Together they bring 74 entity types covering the physical rail network, authorised rolling stock, certification records, and individual vehicles into the NGSI-LD / NGSI-v2 world.

era-eradis: 14 candidate models from ERA ERADIS v3.3.3
era-eratv:  8 candidate models from ERA ERATV v3.3.3
era-evr: 14 candidate models from ERA EVR v3.3.3
era-rinf: 38 candidate models from ERA RINF v3.3.3

Candidates are data models translated/transformed from open standards that are only missing that real examples would qualify them to be officially published.  so if you are using any of the candidates you can contact us in this mail adn we will start the process to make them official.

Mind that the ERA data models (from a previous version are already available in the official repository)

Introducing the Smart Data Models Candidates Repository

The Smart Data Models now has a Candidates repository — a low-friction entry point where any existing standard, ontology or open dataset can be translated into the SDM shape (JSON Schema + NGSI-LD context + examples) before anyone commits to the full contribution process. The first entry, translated from the EDINT Infrastructure Ontology, is live today — and it needs your help to move forward.

What is the Candidates repository?

Until now, turning an external standard into a Smart Data Model meant going almost straight to the Incubated repository, where a contributor actively develops the model to the point of passing our full validation suite and preparing it for official publication. That’s the right process once someone is committed to doing the work — but it leaves no room for a lighter step: “here is a standard that looks like it maps well onto NGSI-LD — does anyone want to develop this further?”

The Candidates repository fills that gap. It maps existing standards — ontologies, open data schemas, sector vocabularies — into the SDM format (JSON Schema, JSON-LD context, key-value and NGSI-LD normalized examples) as a starting point, organized standard-first:

standards/<standard-slug>/
  standard-metadata.yaml
  models/<EntityName>/
    schema.json
    context.jsonld
    examples/
      example.json
      example-normalized.jsonld

A candidate is not a finished data model. It’s a proposal: a structurally valid, plausible translation that any user of Smart Data Models can pick up, test against real payloads, comment on, correct, or use as the basis for a full contribution to Incubated.

The first example: the EDINT Infrastructure Ontology

The first entry, edint-infraestructura, translates the EDINT Infrastructure Ontology — a formal OWL vocabulary developed under the EU-funded EDINT (Espacio de Datos para las Infraestructuras Urbanas Inteligentes) project for describing a municipality’s public and private facilities: educational, cultural, social and sports centers, health centers, on-street and off-street parking, and shared-bicycle stations, together with the sensors, observations, management roles and organizations involved in operating them.

Eleven candidate models were generated: Facility, Parking, BikeSharingStation, ObservationPoint, AccessSensor, CountingSensor, AccessObservation, CountingObservation, Violation, ManagementRole and Organization. Every schema is valid JSON Schema (Draft 2020-12) and every example validates against its own schema — but that’s where “finished” stops. In the spirit of being upfront about what a Candidate actually is, here’s exactly where this one stands:

  • Real examples exist for only one of the eleven models. Facility is grounded in a real record from the source ontology’s own example data (a municipal day-care center in Madrid). The other ten — the sensors, observations, violations, roles — have plausible but synthetic example payloads, because the source ontology itself doesn’t ship real instance data for them. This is the single most useful thing the community could contribute right now: real payloads from an actual access-control sensor, a real counting observation, a real management-role record.
  • Two models likely duplicate existing official subjects. Parking and Organization are thin, attribute-free specializations in the source ontology and may well overlap with the existing official dataModel.Parking and dataModel.Organization. This needs a real review before either goes any further.
  • The source ontology is licensed CC-BY-SA-4.0 (ShareAlike), while SDM contributions are conventionally CC-BY-4.0. This hasn’t been resolved and is flagged directly in the candidate’s standard-metadata.yaml for whoever picks this up next.
  • Geolocation needs a second look. The source ontology’s geometry is expressed in a projected coordinate system (not WGS84), so the example coordinates here are illustrative Madrid locations, not verified reprojections of the original data.

None of this makes the candidate useless — quite the opposite. It’s exactly the kind of concrete, checkable list that turns “someone should look into this standard” into a set of small, tractable tasks anyone can pick up.

How a Candidate could become an official data model

There is no formal graduation process yet, and we’d rather propose one in the open than invent it quietly. Here’s a first draft, and we’re actively looking for feedback on it:

  1. Real-world evidence. At least one model in the candidate has real, non-synthetic example payloads from an actual system or dataset using it.
  2. Passes validation. The candidate’s schemas and examples pass the standard SDM test suite with zero blocking failures (the same 12 checks used for Incubated contributions).
  3. No unresolved duplication. Any overlap with an existing official model has been explicitly checked and either resolved (extend the existing model instead) or justified (the candidate covers something genuinely different).
  4. Compatible licensing. The source material’s license allows redistribution under the terms Smart Data Models publishes under.
  5. A short public review. A brief comment period on the candidate’s pull request or a linked issue, open to anyone in the community.

Once a candidate clears these, it moves into the normal Incubated workflow for full development ahead of official publication.

How to get involved

  • Browse the repository: github.com/smart-data-models/Candidates
  • Search all candidates: smart-data-models.github.io/Candidates — a searchable, always up to date index of every candidate standard and model
  • Contribute a real example for any of the eleven EDINT models above — open a pull request against standards/edint-infraestructura/, or an issue if you’re not sure how.
  • Propose a new candidate: if you know a standard, ontology or open dataset that should have an SDM mapping, open an issue describing it, or submit the translated standards/<slug>/ folder directly.
  • Give feedback on the promotion criteria above — they’re a first draft, not a final decision.

Find all published Smart Data Models at smartdatamodels.org/ search menu, and everything currently in development in the Incubated repository.