This is development documentation. Some features are not available in the stable release. View stable documentation.

The NDDataset object

The NDDataset is the main object used by SpectroChemPy.

Like numpy ndarrays, NDDatasets have the capability to be sliced, sorted and subjected to mathematical operations.

But, in addition, NDDatasets may have units, can be masked and each dimension can have coordinates also with units. This makes NDDatasets aware of units compatibility, e.g., for binary operations such as addition or subtraction or during the application of mathematical operations. In addition to or in replacement of numerical data for coordinates, NDDatasets can also have labeled coordinates where labels can be different kinds of objects (strings, datetime objects, numpy ndarrays or other NDDatasets, etc.).

This offers a lot of flexibility in using NDDatasets that, we hope, will be useful for applications. See the :ref:examples-gallery for additional information about such possible applications.

Table of Contents

Introduction

Below (and in the next sections), we try to give an almost complete view of the NDDataset features.

As we will make some reference to the numpy library, we also import it here.

[1]:
import numpy as np

import spectrochempy as scp

We additionally import the three main SpectroChemPy objects that we will use through this tutorial

[2]:
from spectrochempy import Coord
from spectrochempy import CoordSet
from spectrochempy import NDDataset

For a convenient usage of units, we will also directly import ur, the unit registry which contains all available units.

[3]:
from spectrochempy import ur

Multidimensional arrays are defined in Spectrochempy using the NDDataset object.

NDDataset objects mostly behave like numpy’s numpy.ndarray (see for instance numpy quickstart tutorial).

However, unlike raw numpy arrays, the presence of optional properties makes them (hopefully) more appropriate for handling spectroscopic information, which is one of the major objectives of the SpectroChemPy package:

  • mask: Data can be partially masked at will

  • units: Data can have units, allowing units-aware operations

  • CoordSet: Data can have a set of coordinates, one or several per dimension

Additional metadata can also be added to the instances of this class through the meta properties.

1D-Dataset (unidimensional dataset)

In the following example, a minimal 1D dataset is created from a simple list, to which we can add some metadata:

[4]:
d1D = NDDataset(
    [10.0, 20.0, 30.0],
    name="Dataset N1",
    author="Blake and Mortimer",
    description="A dataset from scratch",
    history="creation",
)
d1D
[4]:
NDDataset [Dataset N1] — float64, size: 3
name
:
Dataset N1
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
A dataset from scratch
history
:
2026-09-29 19:18:36+00:00> Creation
Data
title
:
values
:
[ 10 20 30]
size
:
3
[5]:
print(d1D)
NDDataset: [float64] unitless (size: 3)
[6]:
_ = d1D.plot(figsize=(3, 2))
../../../_images/userguide_objects_dataset_dataset_17_1.png

Except few additional metadata such author , created …, there is not much difference with respect to a conventional numpy.array. For example, one can apply numpy ufunc‘s directly to a NDDataset or make basic arithmetic operation with these objects:

[7]:
np.sqrt(d1D)
[7]:
NDDataset [Dataset N1] — float64, size: 3
name
:
Dataset N1
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
A dataset from scratch
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Ufunc sqrt applied.
Data
title
:
values
:
[ 3.162 4.472 5.477]
size
:
3
[8]:
d1D += d1D / 2.0
d1D
[8]:
NDDataset [Dataset N1] — float64, size: 3
name
:
Dataset N1
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
A dataset from scratch
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Inplace binary op: iadd with `Dataset N1`
Data
title
:
values
:
[ 15 30 45]
size
:
3

As seen above, there are some attributes that are automatically added to the dataset:

  • id: This is a unique identifier for the object.

  • name: A short and unique name for the dataset. It will be equal to the automatic id if it is not provided.

  • author: Author determined from the computer name if not provided.

  • created: Date and time of creation.

  • modified: Date and time of modification.

These attributes can be modified by the user, but the id, created and modified attributes are read only.

Some other attributes are defined to describe the data:

  • title: A long name that will be used in plots or in some other operations.

  • history: Timestamped entries for operations that participate in the history contract. It is inspectable but is not a complete provenance record.

  • description: A comment or a description of the object’s purpose or contents.

  • origin: An optional reference to the source of the data.

Here is an example of the use of the NDDataset attributes:

[9]:
d1D.title = "intensity"
d1D.name = "mydataset"
d1D.history = "created from scratch"
d1D.description = "Some experimental measurements"
d1D
[9]:
NDDataset [mydataset] — float64, size: 3
name
:
mydataset
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
Some experimental measurements
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Inplace binary op: iadd with `Dataset N1`
2026-09-29 19:18:37+00:00> Created from scratch
Data
title
:
intensity
values
:
[ 15 30 45]
size
:
3

d1D is a 1D (1-dimensional) dataset with only one dimension.

Some attributes are useful to check this kind of information:

[10]:
d1D.shape  # the shape of 1D contain only one dimension size
[10]:
(3,)
[11]:
d1D.ndim  # the number of dimensions
[11]:
1
[12]:
d1D.dims  # the name of the dimension (it has been automatically attributed)
[12]:
['x']

Note: The names of the dimensions are set automatically. But they can be changed, with the limitation that the name must be a single letter.

[13]:
d1D.dims = ["q"]  # change the list of dim names.
[14]:
d1D.dims
[14]:
['q']

nD-Dataset (multidimensional dataset)

To create a nD NDDataset, we can provide a nD-array like object to the NDDataset instance constructor

[15]:
d3D = NDDataset.random((2, 4, 6))
d3D.title = "energy"
d3D.author = "Someone"
d3D.name = "3D dataset creation"
d3D.history = "created from scratch"
d3D.description = "Some example"
d3D.dims = ["u", "v", "t"]
d3D
[15]:
NDDataset [3D dataset creation] — float64, shape: (u:2, v:4, t:6)
name
:
3D dataset creation
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
Some example
history
:
2026-09-29 19:18:37+00:00> Created using method : random
2026-09-29 19:18:37+00:00> Created from scratch
Data
title
:
energy
values
:
[[[ 0.9886 0.8448 ... 0.8414 0.5501]
[ 0.8859 0.2943 ... 0.5973 0.5868]
[ 0.979 0.3923 ... 0.5537 0.2702]
[ 0.05483 0.7354 ... 0.8223 0.7821]]

[[ 0.8247 0.8018 ... 0.476 0.5678]
[ 0.5273 0.1226 ... 0.9755 0.1824]
[ 0.9657 0.8657 ... 0.6766 0.6812]
[ 0.1652 0.04971 ... 0.7112 0.2672]]]
shape
:
(u:2, v:4, t:6)

We can also add all information in a single statement

[16]:
d3D = NDDataset.random(
    (2, 4, 6),
    dims=["u", "v", "t"],
    title="Energy",
    author="Someone",
    name="3D_dataset",
    history="created from scratch",
    description="a single statement creation example",
)
d3D
[16]:
NDDataset [3D_dataset] — float64, shape: (u:2, v:4, t:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(u:2, v:4, t:6)

Three names are attributed at creation (if they are not provided with the dims attribute, then the names ‘z’,’y’,’x’ are automatically attributed)

[17]:
d3D.dims
[17]:
['u', 'v', 't']
[18]:
d3D.ndim
[18]:
3
[19]:
d3D.shape
[19]:
(2, 4, 6)

About dates and times

The dates and times are stored internally as UTC (Coordinated Universal Time). Timezone information is stored in the timezone attribute. If not set, the default is to use the local timezone, which is probably the most common case.

[20]:
nd = NDDataset()
nd.created
[20]:
'2026-09-29 19:18:37+00:00'

In this case our local timezone has been used by default for the conversion from UTC datetime.

[21]:
nd.local_timezone
[21]:
'Etc/UTC'
[22]:
nd.timezone = "EST"
nd.created
[22]:
'2026-09-29 14:18:37-05:00'

For a list of timezone code (TZ) you can have a look at List_of_tz_database_time_zones.

About the history attribute

An NDDataset owns one authoritative history store. The public history attribute renders it as a readable, timestamped, list-compatible view. The view supports reading, slicing, display, iteration, and comparison, but is read-only: direct mutations such as dataset.history.append(...) raise TypeError. Use annotate(), replace_history(), or clear_history() when a change is intentional.

history_entries provides the structured view as a deep, detached copy. Every entry has exactly four fields:

  • date: the operation timestamp as a Python datetime;

  • operation: a short identifier, or None for a text-only entry;

  • parameters: a detached dictionary of compact values;

  • message: the readable text rendered by history.

Structured and text-only entries normally coexist. Reader messages, user annotations, older histories, and treatments outside the structured coverage may all have operation=None and an empty parameter dictionary.

This autonomous example starts with synthetic spectra and a user annotation, then applies a structured shift and integrates along the same dimension. The integration keeps the complete chronology that it receives.

[23]:
history_demo = NDDataset(
    np.array([[1.0, 2.0, 4.0, 8.0], [2.0, 3.0, 5.0, 9.0]]),
    dims=["y", "x"],
    coordset=[
        Coord([0.0, 1.0], title="sample"),
        Coord(
            [1000.0, 1100.0, 1200.0, 1300.0],
            units="cm^-1",
            title="wavenumber",
        ),
    ],
    name="history_demo",
)
history_demo.annotate("synthetic spectra prepared")
shifted = history_demo.roll(pts=1.9, dim="x")
integrated = shifted.trapezoid(dim="x")
integrated.history
[23]:
['2026-09-29 19:18:37+00:00> Synthetic spectra prepared',
 "2026-09-29 19:18:37+00:00> `roll` shift performed on dimension `x` with parameters: {'pts': 1.9, 'dim': 'x'}",
 '2026-09-29 19:18:37+00:00> Dataset resulting from application of `trapezoid` method']

The final entry names the integration kernel and records the requested dimension, the resolved dimension, and the corresponding axis in its source dataset. Parameter dictionaries are specific to each operation family; there is no universal parameter schema.

[24]:
integration_entry = integrated.history_entries[-1]
integration_entry["operation"], integration_entry["parameters"]
[24]:
('trapezoid', {'requested_dim': 'x', 'resolved_dim': 'x', 'resolved_axis': 1})

A public alias can delegate to a differently named kernel, so structured entries name the kernel that actually ran. When normalization, conversion, rounding, or clamping makes an effective scientific parameter differ from the request, supported families can retain both. Here the discrete shift applies one point even though 1.9 was requested.

[25]:
shift_entry = shifted.history_entries[-1]
(
    shift_entry["parameters"]["scientific_parameters"],
    shift_entry["parameters"]["requested_parameters"],
)
[25]:
({'pts': 1, 'neg': False}, {'pts': 1.9, 'neg': False})

history_entries is detached all the way through nested parameters. Changing the returned copy does not modify the dataset. An explicit list(dataset.history) is likewise an ordinary mutable copy of the readable strings.

[26]:
detached_entries = integrated.history_entries
detached_entries[-1]["parameters"]["resolved_axis"] = 99
integrated.history_entries[-1]["parameters"]["resolved_axis"]
[26]:
1

Explicit history-management methods make destructive intent visible. Assigning a string to history remains a compatibility shorthand for annotate(); assigning a list is a shorthand for replace_history(); assigning None does nothing. Structured mappings, legacy (date, message) pairs, and strings can be supplied to replace_history().

[27]:
editable_history = history_demo.copy()
editable_history.replace_history(["replacement note", "another note"])
editable_history.history
[27]:
['2026-09-29 19:18:38+00:00> Replacement note',
 '2026-09-29 19:18:38+00:00> Another note']
[28]:
editable_history.clear_history()
editable_history.history
[28]:
[]

Both shallow and deep dataset copies receive independent histories, including nested parameter values. Supported native SCP/PSCP persistence preserves structured entries with format version 3; xarray mapping and the NetCDF representation use version 2. Current readers also accept native version 2 and portable version 1 textual histories, retaining their dates and messages without inferring semantics from prose. Files written with the newer structured formats are not guaranteed to be readable by older SpectroChemPy versions. This is deliberately one-way compatibility.

For the forthcoming 1.1.0 release, structured entries cover the following bounded families:

  • dataset selection and slicing, and transposition;

  • out-of-place scalar or NDDataset addition and subtraction;

  • mean(), sum(), std(), and var() when they return an NDDataset;

  • squeeze(), swapdims(), and reshape();

  • point, circular, and Fourier shifts, and zero filling;

  • apodization and phasing;

  • successful mc(), ps(), ht(), and dc() processing paths;

  • trapezoid() and simpson() integration.

The recorded compact parameters depend on the family. Where its contract establishes them, an entry may distinguish requested and effective scientific parameters, the requested and resolved dimension, the axis in the source dataset, and execution mode. These fields do not form a universal replay schema. Not every arithmetic operation, reduction, or transformation is structured. In-place arithmetic remains text-only, and snv() intentionally adds the single text entry SNVTransformer applied in both execution modes. This coverage is present on the development branch for 1.1.0; it is not a claim about an already published stable version.

This is a readable operation log, not exhaustive provenance. Several treatments remain text-only; learned estimator state and every input chronology of a multi-source assembly are not recorded. Histories are linear and separate, and they do not detect arbitrary mutations, construct a provenance graph, or replay computations. The NMR plugin processing trace and estimator-specific logs are separate mechanisms, not NDDataset.history. Exact message wording is intended for people and is not a stable machine-readable contract; inspect structured fields when available.

A short infrared workflow shows how the two views complement each other. We select six spectra in the OH region, calculate a reference from the first three, and subtract that reference from every selected spectrum.

[29]:
spectra = scp.read("irdata/nh4y-activation.spg")
region = spectra[:6, 3700.0:3300.0]
region.name = "OH-region"
reference = region[:3].mean(dim="y")
reference.name = "mean-reference"
corrected = region - reference
corrected.name = "referenced-OH-region"
[30]:
corrected.plot()
[30]:
<Axes: xlabel='wavenumbers $\\mathrm{/\\ \\mathrm{cm}^{-1}}$', ylabel='absorbance $\\mathrm{/\\ \\mathrm{a.u.}}$'>
../../../_images/userguide_objects_dataset_dataset_62_1.png
[31]:
reference.history
[31]:
['2026-09-29 19:18:38+00:00> Imported OMNIC SPG file nh4y-activation.spg',
 '2026-09-29 19:18:38+00:00> Sorted by date',
 '2026-09-29 19:18:38+00:00> Slice extracted: y indices [:6], x coordinates [3700.0:3300.0]',
 '2026-09-29 19:18:38+00:00> Slice extracted: y indices [:3]',
 '2026-09-29 19:18:38+00:00> Mean computed along y']
[32]:
[entry for entry in reference.history_entries if entry["operation"] == "mean"]
[32]:
[{'date': datetime.datetime(2026, 9, 29, 19, 18, 38, tzinfo=datetime.timezone.utc),
  'operation': 'mean',
  'parameters': {'requested_dims': ['y'],
   'resolved_dims': ['y'],
   'all_dimensions': False,
   'keepdims': False},
  'message': 'Mean computed along y'}]
[33]:
corrected.history
[33]:
['2026-09-29 19:18:38+00:00> Imported OMNIC SPG file nh4y-activation.spg',
 '2026-09-29 19:18:38+00:00> Sorted by date',
 '2026-09-29 19:18:38+00:00> Slice extracted: y indices [:6], x coordinates [3700.0:3300.0]',
 '2026-09-29 19:18:38+00:00> Subtracted `mean-reference` from `OH-region`']
[34]:
[
    entry
    for entry in corrected.history_entries
    if entry["operation"] in {"slice", "subtract"}
]
[34]:
[{'date': datetime.datetime(2026, 9, 29, 19, 18, 38, tzinfo=datetime.timezone.utc),
  'operation': 'slice',
  'parameters': {'source_dims': ['y', 'x'],
   'requested': [{'dimension': 'y',
     'kind': 'index_slice',
     'start': None,
     'stop': 6,
     'step': None},
    {'dimension': 'x',
     'kind': 'coordinate_slice',
     'start': 3700.0,
     'stop': 3300.0,
     'step': None}],
   'inplace': False},
  'message': 'Slice extracted: y indices [:6], x coordinates [3700.0:3300.0]'},
 {'date': datetime.datetime(2026, 9, 29, 19, 18, 38, tzinfo=datetime.timezone.utc),
  'operation': 'subtract',
  'parameters': {'sources': [{'role': 'left',
     'kind': 'NDDataset',
     'name': 'OH-region',
     'title': 'absorbance',
     'shape': [6, 416]},
    {'role': 'right',
     'kind': 'NDDataset',
     'name': 'mean-reference',
     'title': 'absorbance',
     'shape': [416]}]},
  'message': 'Subtracted `mean-reference` from `OH-region`'}]

The reference log records that its mean was computed along y. Its separate chronology is not merged into corrected: the final log follows the main dataset and records its import, spectral selection, and subtraction. The subtraction entry instead identifies the two operands in mathematical order. Understanding that the subtracted reference is the mean of the first three spectra therefore requires its own reference.history; the corrected dataset alone does not contain that construction. This keeps both logs linear and readable without presenting either as complete provenance. The SPG import entry displays only nh4y-activation.spg so the log stays portable; the complete source path remains available from spectra.filename and its reader metadata.

Units

One interesting feature of NDDataset is the ability to define units for the internal data.

[35]:
d1D.units = ur.eV  # ur is a registry containing all available units
[36]:
d1D  # note the eV symbol of the units added to the values field below
[36]:
NDDataset [mydataset] — float64, size: 3, eV
name
:
mydataset
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
Some experimental measurements
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Inplace binary op: iadd with `Dataset N1`
2026-09-29 19:18:37+00:00> Created from scratch
Data
title
:
intensity
values
:
[ 15 30 45] eV
size
:
3

This allows to make units-aware calculations:

[37]:
d1D**2  # note the results in eV^2
[37]:
NDDataset [mydataset] — float64, size: 3, eV²
name
:
mydataset
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
Some experimental measurements
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Inplace binary op: iadd with `Dataset N1`
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:38+00:00> Binary operation pow with `2` has been performed
Data
title
:
power(intensity, 2)
values
:
[ 225 900 2025] eV²
size
:
3
[38]:
np.sqrt(d1D)  # note the result in e^0.5
[38]:
NDDataset [mydataset] — float64, size: 3, eV⁰⋅⁵
name
:
mydataset
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
Some experimental measurements
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Inplace binary op: iadd with `Dataset N1`
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:38+00:00> Ufunc sqrt applied.
Data
title
:
sqrt(intensity)
values
:
[ 3.873 5.477 6.708] eV⁰⋅⁵
size
:
3
[39]:
time = 5.0 * ur.second
d1D / time  # here we get results in eV/s
[39]:
NDDataset [mydataset] — float64, size: 3, eV⋅s⁻¹
name
:
mydataset
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
Some experimental measurements
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Inplace binary op: iadd with `Dataset N1`
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:38+00:00> Binary operation truediv with `5.0 s` has been performed
Data
title
:
intensity
values
:
[ 3 6 9] eV⋅s⁻¹
size
:
3

Compatible conversion returns a new dataset by default and leaves d1D unchanged. Conversions such as energy to temperature or wavenumber to wavelength instead use physical equivalence contexts; they are not ordinary same-dimensional rescaling.

[40]:
in_joules = d1D.to("J")
in_joules
[40]:
NDDataset [mydataset] — float64, size: 3, J
name
:
mydataset
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
Some experimental measurements
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Inplace binary op: iadd with `Dataset N1`
2026-09-29 19:18:37+00:00> Created from scratch
Data
title
:
intensity
values
:
[2.403e-18 4.807e-18 7.21e-18] J
size
:
3
[41]:
d1D  # still expressed in eV
[41]:
NDDataset [mydataset] — float64, size: 3, eV
name
:
mydataset
author
:
Blake and Mortimer
created
:
2026-09-29 19:18:36+00:00
description
:
Some experimental measurements
history
:
2026-09-29 19:18:36+00:00> Creation
2026-09-29 19:18:37+00:00> Inplace binary op: iadd with `Dataset N1`
2026-09-29 19:18:37+00:00> Created from scratch
Data
title
:
intensity
values
:
[ 15 30 45] eV
size
:
3

See units and masks guide for verified examples of quantities, compatible and context-enabled conversions, in-place conversion, incompatibility errors, and masks.

Coordinates

The above created d3D dataset has 3 dimensions, but no coordinates for these dimensions. Here arises a big difference with simple numpy arrays:

  • We can add coordinates to each dimension of a NDDataset.

To get the list of all defined coordinates, we can use the coords attribute:

[42]:
d3D.coordset  # no coordinates, so it returns nothing (None)
[43]:
d3D.t  # the same for coordinate  t, v, u which are not yet set

To add coordinates, one way is to set them one by one:

[44]:
d3D.t = (
    Coord.arange(6) * 0.1
)  # we need a sequence of 6 values for `t` dimension (see shape above)
d3D.t.title = "time"
d3D.t.units = ur.seconds
d3D.coordset  # now return a list of coordinates
[44]:
CoordSet — t:time, u, v
Dimension `t`
size
:
6
title
:
time
coordinates
:
[ 0 0.1 0.2 0.3 0.4 0.5] s
Dimension `u`
title
:
coordinates
:
Undefined
Dimension `v`
title
:
coordinates
:
Undefined
[45]:
d3D.t
[45]:
Coord [t:time] — float64, size: 6, s
size
:
6
title
:
time
coordinates
:
[ 0 0.1 0.2 0.3 0.4 0.5] s
[46]:
d3D.coordset("t")  # Alternative way to get a given coordinates
[46]:
Coord [t:time] — float64, size: 6, s
size
:
6
title
:
time
coordinates
:
[ 0 0.1 0.2 0.3 0.4 0.5] s
[47]:
d3D["t"]  # another alternative way to get a given coordinates
[47]:
Coord [t:time] — float64, size: 6, s
size
:
6
title
:
time
coordinates
:
[ 0 0.1 0.2 0.3 0.4 0.5] s

The two other coordinates u and v are still undefined

[48]:
d3D.u, d3D.v
[48]:
(Coord: empty, Coord: empty)

When the dataset is printed, only the information for the existing coordinates is given.

[49]:
d3D
[49]:
NDDataset [3D_dataset] — float64, shape: (u:2, v:4, t:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(u:2, v:4, t:6)
Dimension `t`
size
:
6
title
:
time
coordinates
:
[ 0 0.1 0.2 0.3 0.4 0.5] s
Dimension `u`
title
:
coordinates
:
Undefined
Dimension `v`
title
:
coordinates
:
Undefined

Programmatically, we can use the attribute is_empty or has_data to check this

[50]:
d3D.v.has_data, d3D.v.is_empty
[50]:
(False, True)

An error is raised when a coordinate doesn’t exist

[51]:
try:
    d3D.x
except Exception as e:
    scp.error_(Exception, e)
 ERROR | Exception:

In some case it can also be useful to get a coordinate from its title instead of its name (the limitation is that if several coordinates have the same title, then only the first ones that is found in the coordinate list, will be returned - this can be ambiguous)

[52]:
d3D["time"]
[52]:
Coord [t:time] — float64, size: 6, s
size
:
6
title
:
time
coordinates
:
[ 0 0.1 0.2 0.3 0.4 0.5] s
[53]:
d3D.time
[53]:
Coord [t:time] — float64, size: 6, s
size
:
6
title
:
time
coordinates
:
[ 0 0.1 0.2 0.3 0.4 0.5] s

Labels

It is possible to use labels instead of numerical coordinates. Labels are sequences of objects. The length of the sequence must be equal to the size of the dimension.

[54]:
tags = list("ab")
d3D.u.title = "some tags"
d3D.u.labels = tags
d3D
[54]:
NDDataset [3D_dataset] — float64, shape: (u:2, v:4, t:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(u:2, v:4, t:6)
Dimension `t`
size
:
6
title
:
time
coordinates
:
[ 0 0.1 0.2 0.3 0.4 0.5] s
Dimension `u`
size
:
2
title
:
some tags
labels
:
[ a b]
Dimension `v`
title
:
coordinates
:
Undefined

or more complex objects.

For instance here we use datetime.timedelta objects:

[55]:
from datetime import timedelta

start = timedelta(0)
times = [start + timedelta(seconds=x * 60) for x in range(6)]
d3D.t = None
d3D.t.labels = times
d3D.t.title = "time"
d3D
[55]:
NDDataset [3D_dataset] — float64, shape: (u:2, v:4, t:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(u:2, v:4, t:6)
Dimension `t`
size
:
6
title
:
time
labels
:
[ 0:00:00 0:01:00 0:02:00 0:03:00 0:04:00 0:05:00]
Dimension `u`
size
:
2
title
:
some tags
labels
:
[ a b]
Dimension `v`
title
:
coordinates
:
Undefined

In this case, getting a coordinate that doesn’t possess numerical data but labels, will return the labels

[56]:
d3D.time
[56]:
Coord [t:time] — size: 6
size
:
6
title
:
time
labels
:
[ 0:00:00 0:01:00 0:02:00 0:03:00 0:04:00 0:05:00]

More insight on coordinates

Sharing coordinates between dimensions

Sometimes it is not necessary to have different coordinates for each axis. Some can be shared between axes.

For example, if we have a square matrix with the same coordinate in the two dimensions, the second dimension can refer to the first. Here we create a square 2D dataset, using the diag method:

[57]:
nd = NDDataset.diag((3, 3, 2.5))
nd
[57]:
NDDataset — float64, shape: (y:3, x:3)
name
:
NDDataset_932173c0
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:18:38+00:00
history
:
2026-09-29 19:18:38+00:00> Created using method : diag
Data
title
:
values
:
[[ 3 0 0]
[ 0 3 0]
[ 0 0 2.5]]
shape
:
(y:3, x:3)

and then we add the same coordinate for both dimensions

[58]:
coordx = Coord.arange(3)
nd.set_coordset(x=coordx, y="x")
nd
[58]:
NDDataset — float64, shape: (y:3, x:3)
name
:
NDDataset_932173c0
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:18:38+00:00
history
:
2026-09-29 19:18:38+00:00> Created using method : diag
Data
title
:
values
:
[[ 3 0 0]
[ 0 3 0]
[ 0 0 2.5]]
shape
:
(y:3, x:3)
Dimension `x`=y
size
:
3
title
:
coordinates
:
[ 0 1 2]

Setting coordinates using set_coordset

Let’s create 3 Coord objects to be used as coordinates for the 3 dimensions of the previous d3D dataset.

[59]:
d3D.dims = ["t", "v", "u"]
s0, s1, s2 = d3D.shape
coord0 = Coord.linspace(10.0, 100.0, s0, units="m", title="distance")
coord1 = Coord.linspace(20.0, 25.0, s1, units="K", title="temperature")
coord2 = Coord.linspace(0.0, 1000.0, s2, units="hour", title="elapsed time")

Syntax 1

[60]:
d3D.set_coordset(u=coord2, v=coord1, t=coord0)
d3D
[60]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

Syntax 2

[61]:
d3D.set_coordset({"u": coord2, "v": coord1, "t": coord0})
d3D
[61]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

Adding several coordinates to a single dimension

We can add several coordinates to the same dimension

[62]:
coord1b = Coord([1, 2, 3, 4], units="millitesla", title="magnetic field")
[63]:
d3D.set_coordset(u=coord2, v=[coord1, coord1b], t=coord0)
d3D
[63]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

We can retrieve the various coordinates for a single dimension easily:

[64]:
d3D.v_1
[64]:
Coord [_1:magnetic field] — float64, size: 4, mT
size
:
4
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT

Math operations on coordinates

Arithmetic operations can be performed on single coordinates:

[65]:
d3D.u = d3D.u * 2
d3D.u
[65]:
Coord [u:elapsed time] — float64, size: 6, h
size
:
6
title
:
elapsed time
coordinates
:
[ 0 400 800 1200 1600 2000] h

The ufunc numpy functions can also be applied, and will affect both the magnitude and the units of the coordinates:

[66]:
d3D.u = 1.5 + np.sqrt(d3D.u)
d3D.u
[66]:
Coord [u:elapsed time] — float64, size: 6, h⁰⋅⁵
size
:
6
title
:
elapsed time
coordinates
:
[ 1.5 21.5 29.78 36.14 41.5 46.22] h⁰⋅⁵

A particularly frequent use case is to subtract the initial value from a coordinate. This can be done directly with the - operator:

[67]:
d3D.u = d3D.u - d3D.u[0]
d3D.u
[67]:
Coord [u:elapsed time] — float64, size: 6, h⁰⋅⁵
size
:
6
title
:
elapsed time
coordinates
:
[ 0 20 28.28 34.64 40 44.72] h⁰⋅⁵

The operations above will generally not work on multiple coordinates, and will raise an error if attempted:

[68]:
try:
    d3D.v = d3D.v - 1.5
except NotImplementedError as e:
    scp.error_(NotImplementedError, e)
 ERROR | NotImplementedError: Subtraction f a CoordSet with an object of type <class 'float'> is not implemented yet

Only subtraction between multiple coordinates is allowed, and will return a new CoordSet where each coordinate has been subtracted:

[69]:
d3D.v = d3D.v - d3D.v[0]
d3D.v
[69]:
CoordSet [v] — _1:magnetic field, _2:temperature
Coord `_1`
size
:
4
title
:
magnetic field
coordinates
:
[ 0 1 2 3] mT
Coord `_2`
size
:
4
title
:
temperature
coordinates
:
[ 0 1.667 3.333 5] K

It is always possible to carry out operations on a given coordinate of a CoordSet. This must be done by accessing the coordinate by its name, e.g. 'temperature' or '_2' for the second coordinate of the v dimension:

[70]:
d3D.v["_2"] = d3D.v["_2"] + 5.0
d3D.v
[70]:
CoordSet [v] — _1:magnetic field, _2:temperature
Coord `_1`
size
:
4
title
:
magnetic field
coordinates
:
[ 0 1 2 3] mT
Coord `_2`
size
:
4
title
:
temperature
coordinates
:
[ 5 6.667 8.333 10] K

Summary of the coordinate setting syntax

Some additional information about coordinate setting syntax

A. First syntax (probably the safer because the name of the dimension is specified, so this is less prone to errors!)

[71]:
d3D.set_coordset(u=coord2, v=[coord1, coord1b], t=coord0)
# or equivalent
d3D.set_coordset(u=coord2, v=CoordSet(coord1, coord1b), t=coord0)
d3D
[71]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

B. Second syntax assuming the coordinates are given in the order of the dimensions.

Remember that we can check this order using the dims attribute of a NDDataset

[72]:
d3D.dims
[72]:
['t', 'v', 'u']
[73]:
d3D.set_coordset((coord0, [coord1, coord1b], coord2))
# or equivalent
d3D.set_coordset(coord0, CoordSet(coord1, coord1b), coord2)
d3D
[73]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

C. Third syntax (from a dictionary)

[74]:
d3D.set_coordset({"t": coord0, "u": coord2, "v": [coord1, coord1b]})
d3D
[74]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

D. It is also possible to use directly the CoordSet property

[75]:
d3D.coordset = coord0, [coord1, coord1b], coord2
d3D
[75]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K
(_2)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
[76]:
d3D.coordset = {"t": coord0, "u": coord2, "v": [coord1, coord1b]}
d3D
[76]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K
[77]:
d3D.coordset = CoordSet(t=coord0, u=coord2, v=[coord1, coord1b])
d3D
[77]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

WARNING

Do not use lists for setting multiple coordinates across different dimensions! Use tuples instead.

Lists have a special meaning in SpectroChemPy - they’re used to set multiple coordinates for the same dimension.

This raises an error (lists have another meaning: they’re used to set multiple coordinates for the same dimension, as shown in examples A and B above):

[78]:
try:
    d3D.coordset = [coord0, coord1, coord2]
except ValueError:
    scp.error_(
        ValueError,
        "Coordinates must be of the same size for a dimension with multiple coordinates",
    )
 ERROR | ValueError: Coordinates must be of the same size for a dimension with multiple coordinates

This works: it uses a tuple (), not a list []

[79]:
d3D.coordset = (
    coord0,
    coord1,
    coord2,
)  # equivalent to d3D.coordset = coord0, coord1, coord2
d3D
[79]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

E. Setting the coordinates individually

Either a single coordinate

[80]:
d3D.u = coord2
d3D
[80]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

or multiple coordinates for a single dimension

[81]:
d3D.v = [coord1, coord1b]
d3D
[81]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K
(_2)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT

or using a CoordSet object.

[82]:
d3D.v = CoordSet(coord1, coord1b)
d3D
[82]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
Someone
created
:
2026-09-29 19:18:37+00:00
description
:
a single statement creation example
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

Methods to create NDDataset

There are many ways to create NDDataset objects.

Let’s first create 2 coordinate objects, for which we can define labels and units. Note the use of the function linspace to generate the data.

[83]:
c0 = Coord.linspace(
    start=4000.0, stop=1000.0, num=5, labels=None, units="cm^-1", title="wavenumber"
)
[84]:
c1 = Coord.linspace(
    10.0, 40.0, 3, labels=["Cold", "RT", "Hot"], units="K", title="temperature"
)

The full coordset will be the following

[85]:
cs = CoordSet(c0, c1)
cs
[85]:
CoordSet — x:temperature, y:wavenumber
Dimension `x`
size
:
3
title
:
temperature
coordinates
:
[ 10 25 40] K
labels
:
[ Cold RT Hot]
Dimension `y`
size
:
5
title
:
wavenumber
coordinates
:
[ 4000 3250 2500 1750 1000] cm⁻¹

Now we will generate the full dataset using the fromfunction method. All needed information is passed as parameters to the NDDataset constructor.

Create a dataset from a function

[86]:
def func(x, y, extra):
    return x * y / extra
[87]:
ds = NDDataset.fromfunction(
    func,
    extra=100 * ur.cm**-1,  # extra arguments passed to the function
    coordset=cs,
    name="mydataset",
    title="absorbance",
    units=None,
)  # when None, units will be determined from the function results

ds.description = """Dataset example created for this tutorial.
It's a 2-D dataset"""

ds.author = "Blake & Mortimer"
ds
[87]:
NDDataset [mydataset] — float64, shape: (y:5, x:3), K
name
:
mydataset
author
:
Blake & Mortimer
created
:
2026-09-29 19:18:39+00:00
description
:
Dataset example created for this tutorial.
It's a 2-D dataset
history
:
2026-09-29 19:18:39+00:00> Created using method : fromfunction
Data
title
:
absorbance
values
:
[[ 400 1000 1600]
[ 325 812.5 1300]
...
[ 175 437.5 700]
[ 100 250 400]] K
shape
:
(y:5, x:3)
Dimension `x`
size
:
3
title
:
temperature
coordinates
:
[ 10 25 40] K
labels
:
[ Cold RT Hot]
Dimension `y`
size
:
5
title
:
wavenumber
coordinates
:
[ 4000 3250 2500 1750 1000] cm⁻¹

Using numpy-like constructors of NDDatasets

[88]:
dz = NDDataset.zeros(
    (5, 3), coordset=cs, units="meters", title="Datasets with only zeros"
)
[89]:
do = NDDataset.ones(
    (5, 3), coordset=cs, units="kilograms", title="Datasets with only ones"
)
[90]:
df = NDDataset.full(
    (5, 3), fill_value=1.25, coordset=cs, units="radians", title="with only float=1.25"
)
df
[90]:
NDDataset — float64, shape: (y:5, x:3), rad
name
:
NDDataset_93a886ae
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:18:39+00:00
history
:
2026-09-29 19:18:39+00:00> Created using method : full
Data
title
:
with only float=1.25
values
:
[[ 1.25 1.25 1.25]
[ 1.25 1.25 1.25]
...
[ 1.25 1.25 1.25]
[ 1.25 1.25 1.25]] rad
shape
:
(y:5, x:3)
Dimension `x`
size
:
3
title
:
temperature
coordinates
:
[ 10 25 40] K
labels
:
[ Cold RT Hot]
Dimension `y`
size
:
5
title
:
wavenumber
coordinates
:
[ 4000 3250 2500 1750 1000] cm⁻¹

As with numpy, it is also possible to take another dataset as a template:

[91]:
df = NDDataset.full_like(d3D, dtype="int", fill_value=2)
df
[91]:
NDDataset [3D_dataset] — float64, shape: (t:2, v:4, u:6)
name
:
3D_dataset
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:18:39+00:00
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
2026-09-29 19:18:39+00:00> Created using method : full_like
Data
title
:
Energy
values
:
[[[ 2 2 ... 2 2]
[ 2 2 ... 2 2]
[ 2 2 ... 2 2]
[ 2 2 ... 2 2]]

[[ 2 2 ... 2 2]
[ 2 2 ... 2 2]
[ 2 2 ... 2 2]
[ 2 2 ... 2 2]]]
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K
[92]:
nd = NDDataset.diag((3, 3, 2.5))
nd
[92]:
NDDataset — float64, shape: (y:3, x:3)
name
:
NDDataset_93a886f6
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:18:39+00:00
history
:
2026-09-29 19:18:39+00:00> Created using method : diag
Data
title
:
values
:
[[ 3 0 0]
[ 0 3 0]
[ 0 0 2.5]]
shape
:
(y:3, x:3)

Copying existing NDDataset

To copy an existing dataset, this is as simple as:

[93]:
d3D_copy = d3D.copy()

or alternatively:

[94]:
d3D_copy_alt = d3D[:]

Finally, it is also possible to initialize a dataset using an existing one:

[95]:
d3Dduplicate = NDDataset(d3D, name=f"duplicate of {d3D.name}", units="absorbance")
d3Dduplicate
[95]:
NDDataset [duplicate of 3D_dataset] — float64, shape: (t:2, v:4, u:6), a.u.
name
:
duplicate of 3D_dataset
author
:
runner@runnervmtr4k5
created
:
2026-09-29 19:18:39+00:00
history
:
2026-09-29 19:18:37+00:00> Created from scratch
2026-09-29 19:18:37+00:00> Created using method : random
Data
title
:
Energy
values
:
[[[ 0.5543 0.902 ... 0.6938 0.2746]
[ 0.04433 0.2064 ... 0.1952 0.9148]
[ 0.1088 0.2259 ... 0.5132 0.9216]
[ 0.2775 0.4854 ... 0.3608 0.5566]]

[[ 0.1416 0.7667 ... 0.1449 0.8738]
[ 0.2606 0.693 ... 0.9218 0.3677]
[ 0.4676 0.8556 ... 0.9343 0.1164]
[ 0.841 0.2715 ... 0.513 0.4529]]] a.u.
shape
:
(t:2, v:4, u:6)
Dimension `t`
size
:
2
title
:
distance
coordinates
:
[ 10 100] m
Dimension `u`
size
:
6
title
:
elapsed time
coordinates
:
[ 0 200 400 600 800 1000] h
Dimension `v`
size
:
4
(_1)
title
:
magnetic field
coordinates
:
[ 1 2 3 4] mT
(_2)
title
:
temperature
coordinates
:
[ 20 21.67 23.33 25] K

Importing from external datasets

NDDatasets can be created from the importation of external data.

A test data folder contains some sample data for experimenting with features of datasets.

[96]:
# let check if this directory exists and display its actual content:
datadir = scp.preferences.datadir
if datadir.exists():
    print(datadir.name)
testdata

Let’s load grouped IR spectra acquired using OMNIC:

[97]:
nd = scp.read_omnic(datadir / "irdata/nh4y-activation.spg")
scp.preferences.reset()
_ = nd.plot()
../../../_images/userguide_objects_dataset_dataset_187_0.png

Even if we do not specify the datadir, the application first looks in the default directory.

Now, lets load a NMR dataset (in the Bruker format).

Requires the official spectrochempy-nmr plugin. Install with: python -m pip install "spectrochempy[nmr]".

[98]:
path = datadir / "nmrdata" / "bruker" / "tests" / "nmr" / "topspin_1d"

# load the data directly (no need to create the dataset first)
nd2 = scp.nmr.read(path, expno=1, remove_digital_filter=True)

# view it...
nd2.x.to("s")

ax = nd2.plot()
../../../_images/userguide_objects_dataset_dataset_190_0.png