Skip to content

Soundscape

Bridge to ambiscape: soundscape features on the MGT time base.

MGT owns pixels, ambiscape owns samples; this adapter is the one crossing point. It runs (or reuses) ambiscape's cached feature extraction and returns the 1 Hz series as an MgFeatures container whose metadata carries the absolute start time, so motion and audio series join on the wall clock. Requires pip install "musicalgestures[soundscape]".

A video is a valid input. The original adapter took an ambiscape session folder --- several WAVs on one clock --- which is right when somebody recorded a place on purpose and wrong when what you have is a video, which is what this toolbox is for. ambiscape opens a single file as a one-take session, so a video needs only its audio pulling out first.

The part with a right answer is which of the three a path is, and audio_source_for is that decision on its own so it can be tested without ffmpeg or ambiscape present. Getting it wrong silently is how a video ends up analysed as a folder containing one thing.

features_from_ambiscape

features_from_ambiscape(F)

ambiscape's feature dict as named 1 Hz series.

Level, the shape of the spectrum, how tonal or noisy it is, and where the energy sits by octave band --- which is a soundscape vocabulary, where level alone is only loud or quiet.

Parameters:

Name Type Description Default
F

the mapping ambiscape's load_features returns.

required

Returns:

Name Type Description
dict dict

name to series, all the same length, all on the 1 Hz grid.

Raises:

Type Description
KeyError

if rms_w is absent. Level is the one thing every recording has, and a feature set without it is not one this can read.

Source code in musicalgestures/_soundscape.py
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
def features_from_ambiscape(F) -> dict:
    """ambiscape's feature dict as named 1 Hz series.

    Level, the shape of the spectrum, how tonal or noisy it is, and where the energy sits
    by octave band --- which is a soundscape vocabulary, where level alone is only loud or
    quiet.

    Args:
        F: the mapping ambiscape's `load_features` returns.

    Returns:
        dict: name to series, all the same length, all on the 1 Hz grid.

    Raises:
        KeyError: if `rms_w` is absent. Level is the one thing every recording has, and a
            feature set without it is not one this can read.
    """
    rms = np.asarray(F["rms_w"], float)
    #: +1e-12 before the log. A silent block is a real state --- a room with nobody in it
    #: --- and log10(0) is -inf, which poisons every mean and every plot downstream.
    out = {"aud_level_db": 20 * np.log10(rms + 1e-12)}

    for src_name, dst_name in _SCALAR_FEATURES.items():
        if src_name in F:
            v = np.asarray(F[src_name], float)
            if v.shape[:1] == rms.shape[:1]:
                out[dst_name] = v

    if "oct_pow" in F:
        oct_pow = np.asarray(F["oct_pow"], float)
        if oct_pow.ndim == 2 and oct_pow.shape[0] == len(rms):
            for i in range(oct_pow.shape[1]):
                out[f"aud_oct{i:02d}_db"] = 10 * np.log10(oct_pow[:, i] + 1e-12)
    return out

audio_source_for

audio_source_for(source, audio=None)

What kind of input this is, and what to hand ambiscape.

Parameters:

Name Type Description Default
source

a directory (an ambiscape session), an audio file, or a video.

required
audio

an audio file to use instead of extracting one. The caller often already has it, and extracting it again is waste.

None

Returns:

Name Type Description
tuple

(kind, path) where kind is "session", "file" or "extract".

For "extract" the path is where the audio should be written; it does not

exist yet, and it is deliberately not the source's own name with the suffix

swapped, so a source that is already an audio file cannot be overwritten by its

own extraction.

Raises:

Type Description
FileNotFoundError

if source does not exist.

ValueError

if it is a file of a kind this cannot open.

Source code in musicalgestures/_soundscape.py
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
def audio_source_for(source, audio=None):
    """What kind of input this is, and what to hand ambiscape.

    Args:
        source: a directory (an ambiscape session), an audio file, or a video.
        audio: an audio file to use instead of extracting one. The caller often already
            has it, and extracting it again is waste.

    Returns:
        tuple: ``(kind, path)`` where kind is ``"session"``, ``"file"`` or ``"extract"``.
        For ``"extract"`` the path is where the audio should be written; it does not
        exist yet, and it is deliberately not the source's own name with the suffix
        swapped, so a source that is already an audio file cannot be overwritten by its
        own extraction.

    Raises:
        FileNotFoundError: if `source` does not exist.
        ValueError: if it is a file of a kind this cannot open.
    """
    from pathlib import Path as _P

    src = _P(source)
    if not src.exists():
        raise FileNotFoundError(f"{src} does not exist")
    if audio is not None:
        return "file", _P(audio)
    if src.is_dir():
        return "session", src
    suffix = src.suffix.lower()
    if suffix in _AUDIO_SUFFIXES:
        return "file", src
    if suffix in _VIDEO_SUFFIXES:
        #: The container goes in the name. Two recordings called clip.mov and clip.mp4
        #: beside each other would otherwise both extract to clip.wav, and the second
        #: run would silently analyse the first one's audio.
        return "extract", src.with_name(f"{src.stem}_{suffix[1:]}_soundscape.wav")
    raise ValueError(
        f"{src.name}: cannot take soundscape features from a {suffix or 'suffixless'} "
        f"file. Give a video, an audio file, or an ambiscape session folder.")

soundscape_features

soundscape_features(source, features_dir=None, audio=None, sr=48000)

Soundscape features as an MgFeatures (1 Hz, wall-clocked).

Parameters:

Name Type Description Default
source

an ambiscape session folder, an audio file, or a video, whose audio is extracted once and reused.

required
features_dir

cache directory for ambiscape's .npz features (default: <source>/analysis/features for a folder, or beside the audio).

None
audio

an audio file to use instead of extracting one from source.

None
sr int

Sample rate for the extraction. Defaults to 48000, which is what ambiscape's band measures expect; downsampling first would move the noise floor it reports.

48000
Source code in musicalgestures/_soundscape.py
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
def soundscape_features(source, features_dir=None, audio=None,
                        sr: int = 48000) -> MgFeatures:
    """Soundscape features as an MgFeatures (1 Hz, wall-clocked).

    Args:
        source: an ambiscape session folder, an audio file, or **a video**, whose audio is
            extracted once and reused.
        features_dir: cache directory for ambiscape's .npz features
            (default: ``<source>/analysis/features`` for a folder, or beside the audio).
        audio: an audio file to use instead of extracting one from `source`.
        sr (int): Sample rate for the extraction. Defaults to 48000, which is what
            ambiscape's band measures expect; downsampling first would move the noise
            floor it reports.
    """
    import subprocess
    try:
        import ambiscape as asc
        from ambiscape import features as afeat
    except ImportError as e:
        raise ImportError(
            "ambiscape is required: pip install "
            "'musicalgestures[soundscape]'") from e

    kind, path = audio_source_for(source, audio=audio)
    if kind == "extract" and not path.exists():
        #: Once. A second call finds the file and skips this, which matters when the
        #: source is a two-hour recording on an external drive.
        subprocess.run(["ffmpeg", "-v", "error", "-y", "-i", str(Path(source)),
                        "-vn", "-ac", "1", "-ar", str(sr), str(path)], check=True)

    if kind == "session":
        sess = asc.open_session(path)
        out = Path(features_dir) if features_dir else path / "analysis" / "features"
    else:
        sess = asc.open_recording(str(path))
        out = Path(features_dir) if features_dir else path.parent / "features"
    npz = afeat.extract_session(sess, out, verbose=False)
    F = afeat.load_features(sorted(Path(p) for p in npz))

    day0_midnight = dt.datetime.combine(sess.day0, dt.time())
    return MgFeatures(
        features_from_ambiscape(F),
        times=np.asarray(F["t"], float),
        sr=1.0,
        source=str(path),
        metadata={"start_datetime": day0_midnight.isoformat(),
                  "tool": "ambiscape"},
    )

merge_into_summary

merge_into_summary(features, summary_json, prefix='mot_')

Fold feature medians/IQRs into an analysis summary.json.

The mirror of ambiscape's vision --merge (which uses vis_): each feature contributes <prefix><name>_median and <prefix><name>_iqr so one summary file describes the whole audio–video session. Existing keys are preserved.

Source code in musicalgestures/_soundscape.py
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
def merge_into_summary(features: MgFeatures, summary_json,
                       prefix: str = "mot_"):
    """Fold feature medians/IQRs into an analysis summary.json.

    The mirror of ambiscape's ``vision --merge`` (which uses ``vis_``):
    each feature contributes ``<prefix><name>_median`` and
    ``<prefix><name>_iqr`` so one summary file describes the whole
    audio–video session. Existing keys are preserved.
    """
    import json

    summary_json = Path(summary_json)
    doc = json.loads(summary_json.read_text()) \
        if summary_json.exists() else {}
    for name in features.feature_names:
        x = np.asarray(features[name], dtype=float)
        q25, q50, q75 = np.nanpercentile(x, [25, 50, 75])
        doc[f"{prefix}{name}_median"] = round(float(q50), 4)
        doc[f"{prefix}{name}_iqr"] = round(float(q75 - q25), 4)
    summary_json.write_text(json.dumps(doc, indent=2))
    return summary_json