TACET: Context-Appropriate Acoustic-Social
Navigation for Quadrupeds

Anonymous Authors
Submitted to IEEE ICRA 2027

Overview

A legged robot's locomotion noise is a social variable. TACET infers the social context of the scene and emits one compact behavior token, ⟨gait, speed, social_cost⟩, that sets both where the robot goes and how loudly it walks.

9.3 dBA
quieter than the default gait, at matched speed
100%
personal-space compliance, every scenario
≤2.9 dBA
added at the listener's own position
7
real-world encounters on a Unitree Go2

Abstract

Quadruped robots entering hospitals, care homes, and quiet offices must be context-appropriate not only in where they move but in how loudly they move: a legged robot's locomotion noise, dominated by foot–ground impacts, is itself a social variable. Prior social navigation respects human space but treats the robot as acoustically uniform, while quiet-locomotion methods reduce noise to an operator-specified, context-blind level. We present TACET, a context-appropriate acoustic-social navigation method that infers social context from the robot's egocentric view and decides both where it walks and how loudly, coupling a slow fine-tuned vision-language reasoner to a fast reactive controller through a single compact behavior token, ⟨gait, speed, social_cost⟩. The same token conditions both a social costmap (where to go) and a quiet locomotion policy (how loudly to move), while a structured out-of-view memory keeps recently seen people in the reasoner's context after they leave the camera view. On a real quadruped, context-conditioned locomotion lowers locomotion noise by up to 9.3 dBA at matched speed, and across our scenarios the full method keeps personal-space compliance at 100% with low acoustic intrusion (≤2.9 dBA), jointly improving spatial and acoustic performance in the evaluated scenarios.

A Dual-Process Architecture

01

Read the social scene

A fine-tuned vision-language model reads the egocentric view and emits one behavior token at roughly 0.25 Hz — slow enough to reason, far outside the safety loop.

02

One token, two mechanisms

The same token scales an asymmetric proxemic costmap and conditions a quiet locomotion policy, so room and quietness are allocated by a single inferred signal.

03

Remember who left the frame

A structured out-of-view memory carries recently seen people across field-of-view loss, each tagged with an age-decayed confidence.

Overview of TACET: person tracker, social memory, VLM reasoner, social costmap, path planner, velocity control, locomotion policy
The semantic reasoner (person tracker → social memory → VLM reasoner) emits the behavior token; the reactive controller (social costmap → path planner → velocity controller → locomotion policy) consumes it to produce socially aware, acoustically controlled motion. The token is the sole interface: the reasoner never outputs motor commands, and the controller never performs semantic reasoning.

The semantic reasoner

A fine-tuned Qwen2.5-VL-7B maps the egocentric RGB image, a natural-language instruction, and a structured social memory to a short chain-of-thought followed by the behavior token, with gait ∈ {quiet, normal, urgent}, speed ∈ {low, normal, high}, and social_cost a scalar amplitude at a resolution of 0.5. Two categorical fields and one scalar make the output directly parseable, in contrast to free-form language or continuous embeddings.

Supervision comes from auto-labeling: a large VLM produces structured ⟨thought, action token⟩ labels on 3,950 egocentric frames from MuSoHu, JRDB-Social, SCAND and InCrowd-VI, every one human-verified, combined with 18,042 COCO images auto-labeled into the same format and 909 SocialNav-SUB frames remapped onto the token format, for 34,751 training records over 22,901 distinct frames. On a held-out benchmark of 100 human-labeled SCAND frames, exact agreement rises from 30% / 26% / 14% for the base model to 95% / 96% / 69% after QLoRA fine-tuning. A query takes 2.6 s at the median, so tokens are refreshed at roughly 0.25 Hz and cached — slow reasoning never enters the safety loop.

Out-of-view social memory

A narrow camera field of view means people the robot was just reacting to leave the frame while remaining well inside its local planning neighborhood. People are detected with YOLOv8n-pose, localized by fusing depth with the point cloud, and maintained in the odometry frame under a hold / coast / refresh policy: a moving person is coasted at constant velocity, a stationary person held until a time-to-live expires. Each retained person becomes a structured record whose confidence starts from source quality and decays with observation age. Records enter the prompt as reasoning context, never as a rule-based override.

Reactive control — costmap and gait, from the same token

Each person induces an anisotropic Gaussian in their own frame, written into one layer of a layered costmap. What differs from prior proxemic costmaps is not the shape but where its parameters come from: in TACET the amplitude and width follow the emitted token, so the same decision that sets how quietly the robot walks also sets how much room it leaves. Collisions stay in a separate hard obstacle layer, so social cost modulates preference without compromising geometric safety.

Because quadruped locomotion noise is dominated by foot–ground impacts, whose acoustic energy grows with foot velocity at contact, TACET reduces perceived noise through control rather than hardware. The gait and speed tokens enter the policy as physical quantities rather than a one-hot code: gait selects a contact schedule, a four-beat crawl for quiet and a trot with decreasing duty factor for normal and urgent, and speed sets its period. One network therefore covers every gait–speed pair the reasoner emits. A token-gated foot-impact penalty tightens as the gait token moves from urgent to quiet at matched speed. Unlike an operator-set quiet factor, quietness is allocated by context rather than preset.

Per-person social fields: circular for a static person, egg-shaped for a moving person, and one enlarged field for an interacting pair
Per-person social fields at the maximum social_cost token: circular (static), egg-shaped elongated along travel (moving), and one enlarged field containing an interacting pair, protecting the shared interaction space.

The Gait as a Noise Source

We first characterize the locomotion policy in isolation, with no people present, which separates the gait's acoustic contribution from that of any detour. Baselines are compared at matched actual speed taken from LiDAR-inertial odometry, so a method cannot appear quieter merely by moving more slowly.

Speed (m/s) Method MNL (dBA) ↓ PNL (dBA) ↓
0.5Default Go268.9 ± 2.872.5 ± 0.7
MUTE (β=1)67.8 ± 2.071.5 ± 0.5
Ours64.1 ± 1.365.9 ± 0.8
0.8Default Go274.4 ± 2.477.2 ± 1.2
MUTE (β=1)71.7 ± 2.974.6 ± 0.7
Ours65.1 ± 2.168.1 ± 1.6
1.1Default Go274.8 ± 1.976.9 ± 1.2
MUTE (β=1)72.3 ± 2.576.6 ± 0.2
Ours70.6 ± 2.475.1 ± 0.6

Speed is the commanded speed. Mean ± standard deviation over ten trials per speed; lower is better.

Ours is the quietest at every matched speed, lowering mean noise level by up to 9.3 dBA over the default controller and 6.6 dBA over MUTE. Standing still with only its cooling fan running, the Go2 already registers about 60 dBA at the same microphone position — the floor these levels should be read against. The margin is smallest at 1.1 m/s, where impact noise dominates and little headroom remains.

Seven Encounters

Nothing in the instruction names a gait, a speed, or an amount of room. The token follows the scene, and the behavior follows the token — from a quiet crawl past someone working to an urgent trot past a siren light.

With a person in the robot's path

Evaluated with the full set of spatial and acoustic social metrics.

Focused person · working indoor

Someone absorbed in a stationary task in a quiet room. The robot must reach a nearby goal without acoustically intruding.

quiet · low · 4.0   3.61 m · 1.6 dBA
Focused person · headphones indoor

The same person, seat, posture and layout. Headphones are the only changed cue — and the gait token changes with them.

normal · low · 4.0   3.62 m · 2.9 dBA
Standing group indoor

A small group converses. The robot must reach a goal beyond them without cutting through the shared interaction space.

quiet · low · 3.5   2.25 m · 2.7 dBA
Moving person outdoor

A person walks through the space, where a higher ambient level masks part of the robot's own noise. The robot yields room as it passes.

quiet · normal · 2.5   1.44 m · 1.7 dBA

With no one in the path

The social metrics do not apply here; what matters is that caution is not spent where it buys nothing.

Empty corridor indoor

No one present, so any detour or slowdown is pure overhead.

normal · normal · 1.5   14.4 s to goal
Open sports field outdoor

A large open space with people active far from the robot's path.

normal · high · 1.0   14.0 s to goal
Emergency indoor

A flashing siren light in a dim corridor. Seeing it, the reasoner decides speed outweighs restraint.

urgent · high · 1.0   11.2 s to goal

Results

PSC is the fraction of the run spent at least 1.2 m from every person; MinDist the closest approach; ΔL the A-weighted level the robot adds at the listener; AID the fraction of the run for which ΔL exceeds 5 dBA.

Scenario Method PSC (%) ↑ MinDist (m) ↑ ΔL (dBA) ↓ AID (%) ↓
Focused (working)RPP95.61.03 ± 0.128.4 ± 0.566.3
SFM92.70.41 ± 0.055.5 ± 0.351.8
MUTE97.40.91 ± 0.146.4 ± 0.361.2
VLM-SN100.01.35 ± 0.065.9 ± 0.258.3
Ours100.03.61 ± 0.171.6 ± 0.20.0
Focused (headph.)RPP95.81.06 ± 0.148.1 ± 0.665.9
SFM92.20.39 ± 0.065.7 ± 0.252.3
MUTE97.10.89 ± 0.156.3 ± 0.360.1
VLM-SN100.01.35 ± 0.045.7 ± 0.257.9
Ours100.03.62 ± 0.192.9 ± 0.226.5
Standing groupRPP92.30.90 ± 0.1912.1 ± 1.1100.0
SFM92.81.02 ± 0.027.8 ± 0.673.3
MUTE93.71.13 ± 0.054.9 ± 0.542.9
VLM-SN90.40.87 ± 0.167.3 ± 0.565.7
Ours100.02.25 ± 0.212.7 ± 0.322.3
Moving personRPP87.70.59 ± 0.057.9 ± 0.783.2
SFM90.20.75 ± 0.045.1 ± 0.551.3
MUTE98.91.21 ± 0.113.8 ± 0.324.7
VLM-SN94.10.69 ± 0.074.1 ± 0.429.4
Ours100.01.44 ± 0.121.7 ± 0.30.0

Means over ten trials, with standard deviations given for MinDist and ΔL. VLM-SN abbreviates VLM-Social-Nav.

Ours is the only method that keeps personal-space compliance at 100% in all four scenarios; VLM-Social-Nav reaches 100% with the focused person but not with the group or the moving person. Ours also keeps a far wider margin at the closest approach, with MinDist never below 1.44 m against at most 1.35 m for any baseline. All runs reached the goal without contacting a person, but not always by the robot's own doing. Three baseline conditions came within arm's reach of a person: SFM in both focused-person conditions, where it repeatedly stalled and scraped the wall (0.41 and 0.39 m), and RPP with the moving person, which stayed clear only because the person stepped aside (0.59 m).

What the reasoner adds

The two focused-person conditions hold the scene fixed and change only the headphones. Methods that consume positions alone behave near-identically in both, as VLM-Social-Nav does (ΔL 5.9 vs. 5.7 dBA). Ours separates them, keeping the same room (MinDist 3.61 vs. 3.62 m) while relaxing the acoustic effort where the listener is less sensitive to it (ΔL 1.6 vs. 2.9 dBA, AID 0.0 vs. 26.5%). With the closest approach unchanged and only the gait token differing, the acoustic difference is consistent with the gait effect measured on the controlled traverse.

What it costs

Scenario Baseline mean (s) Baseline best (s) Ours (s)
Focused (working)31.821.236.8
Focused (headph.)31.822.033.8
Standing group17.213.625.6
Moving person16.714.821.8
Empty corridor16.214.414.4
Sports field18.314.414.0
Emergency15.414.211.2

Mean time to goal over ten trials; the baseline columns summarize RPP, SFM, MUTE and VLM-Social-Nav. Bold marks the fastest entry in each row.

Where people need space or quiet, TACET spends at most 1.5× the mean baseline time. Where they do not, it matches the fastest baseline on an empty corridor and is the fastest method outright on the open field and with the siren light.

Human Study

We ran a blinded video study with 15 observers to complement the objective results. Each observer rated the 20 short clips formed by the four person-facing scenarios and five methods (RPP, SFM, MUTE, VLM-Social-Nav, and TACET), presented in randomized order and without method labels. After every clip, the observer rated Q1–Q4 on a 0–5 agreement scale. All sessions used the same headphones and playback level so that relative footstep-noise differences were preserved.

0: strongly disagree · 1: disagree · 2: somewhat disagree · 3: somewhat agree · 4: agree · 5: strongly agree.

Focused person · working

quiet indoor setting

  1. Q1The robot maintained a safe and comfortable distance from the person.
  2. Q2The robot appropriately adjusted its behavior to the social situation and surroundings.
  3. Q3The robot’s footstep noise was appropriate for this quiet indoor setting.
  4. Q4The robot did not unnecessarily disturb the person’s work.

Focused person · headphones

quiet indoor setting

  1. Q1The robot maintained a safe and comfortable distance from the person.
  2. Q2The robot appropriately adjusted its behavior to the social situation and surroundings.
  3. Q3The robot’s footstep noise was appropriate for this quiet indoor setting.
  4. Q4The robot behaved appropriately around a person wearing headphones.

Standing group

indoor conversation

  1. Q1The robot maintained a safe and comfortable distance from the group.
  2. Q2The robot appropriately adjusted its behavior to the social situation and surroundings.
  3. Q3The robot’s footstep noise was appropriate near the group’s conversation.
  4. Q4The robot avoided interrupting the group’s shared conversational space.

Moving person

outdoor crossing

  1. Q1The robot maintained a safe and comfortable distance from the pedestrian.
  2. Q2The robot appropriately adjusted its behavior to the social situation and surroundings.
  3. Q3The robot’s footstep noise was appropriate for this outdoor setting.
  4. Q4The pedestrian could continue without having to alter their path to avoid the robot.
Per-question results of the blinded video study: mean agreement with standard errors for RPP, SFM, MUTE, VLM-Social-Nav and TACET across the four person-facing scenarios
Per-question results of the blinded video study. Q1: safe and comfortable distance; Q2: contextual consideration; Q3: footstep-noise appropriateness; Q4: scenario-specific etiquette. Bars show mean scores on the 0–5 agreement scale over paired within-observer ratings; error bars show standard errors. Higher is better.

TACET receives the highest agreement in all sixteen scenario–question cells, averaging 4.54 against 2.23 for MUTE, 1.92 for VLM-Social-Nav, 1.07 for SFM and 0.78 for RPP, and its standard errors are the smallest (0.14 on average against 0.29 for VLM-Social-Nav), so observers agreed about it more than about any baseline. The margin narrows only with the moving person, where MUTE reaches 3.33–4.07 — the same scenario in which MUTE is the strongest baseline on the objective metrics, so its quiet gait alone recovers much of what observers reward.

Ablation

One component is removed at a time on two representative person-facing scenarios. Separately, the top block of each scenario keeps the architecture fixed and swaps only the reasoner that emits the token, isolating the contribution of fine-tuning beyond prompt engineering.

Focused person indoor

Clearance is what the social layer buys; quietness is what the trained policy buys. Removing either one leaves the other intact. The clip then swaps the reasoner for a prompted base VLM, which keeps the architecture fixed and loses both.

Moving person outdoor

People pass in and out of the camera view as the robot walks, so this is where the social memory earns its place. The reasoner swap that follows barely changes the route here, yet the robot turns audible above ambient for a third of the run.

Scenario Variant PSC (%) ↑ MinDist (m) ↑ ΔL (dBA) ↓ AID (%) ↓
Focused (working)Qwen2.5-VL-7B100.01.35 ± 0.104.5 ± 0.531.8
InternVL3.5-8B100.01.33 ± 0.124.9 ± 0.435.1
Ours (full)100.03.61 ± 0.171.6 ± 0.20.0
w/o social layer96.51.18 ± 0.114.3 ± 0.329.3
w/o trained policy100.03.47 ± 0.197.4 ± 0.563.3
w/o social memory100.03.54 ± 0.243.9 ± 0.318.7
Moving personQwen2.5-VL-7B100.01.38 ± 0.113.5 ± 0.536.1
InternVL3.5-8B100.01.33 ± 0.123.8 ± 0.739.3
Ours (full)100.01.44 ± 0.121.7 ± 0.30.0
w/o social layer94.70.89 ± 0.123.8 ± 0.337.4
w/o trained policy100.01.39 ± 0.134.2 ± 0.541.5
w/o social memory100.01.42 ± 0.092.6 ± 0.419.3

Means over ten trials, with standard deviations given for MinDist and ΔL. Rows above each rule swap only the reasoner that emits the token; rows below remove one component.

Distance and gait are separable. Removing the trained policy keeps the closest approach within one standard deviation of the full method, yet degrades the acoustic terms the most of any variant, by 5.8 dBA with the focused person and 2.5 dBA with the moving one, supporting an acoustic contribution from the gait policy beyond that of the detour. Removing the social layer costs clearance instead, collapsing MinDist from 3.61 to 1.18 m and from 1.44 to 0.89 m, both inside the 1.2 m zone the full method never enters. Removing social memory barely alters the route, yet the robot becomes audible above ambient for 18.7% and 19.3% of the run: without it, a person who briefly leaves the camera view is no longer carried in the reasoner's context.

Swapping the reasoner degrades both axes. With the focused person, a prompted Qwen2.5-VL-7B keeps only 1.35 m of clearance against 3.61 m for the full method, close to what removing the social layer gives — and raises ΔL to 4.5 dBA with AID at 31.8%; InternVL3.5-8B behaves the same way (1.33 m, 4.9 dBA, 35.1%). With the moving person the route changes far less (1.38 and 1.33 m against 1.44 m), yet the robot is audible above ambient for 36.1% and 39.3% of the run against 0.0% for the full method. Since the architecture is unchanged, the loss follows from the token interface itself rather than from any component.

Implementation

Platform

  • Unitree Go2 quadruped
  • Intel RealSense D435i RGB-D
  • Hesai XT-16 LiDAR
  • Point-LIO odometry
  • Nav2 (Smac Hybrid-A*) + pure pursuit

Reasoner

  • Qwen2.5-VL-7B, 4-bit QLoRA, rank 16
  • One epoch, learning rate 2×10−5
  • 34,751 records / 22,901 frames
  • Median query 2.6 s (p95 3.5 s)
  • Token refresh ≈0.25 Hz, cached
  • NVIDIA RTX A6000 (48 GB)

Locomotion policy

  • MLP [512, 256, 128], ELU
  • 50 Hz, joint PD (Kp=40, Kd=1)
  • Six gait/speed modes, one network
  • Isaac Lab, deployed through rl_sar

Social costmap

Each tracked person contributes a bounded, anisotropic social field that changes the Nav2 path preference; the obstacle layer remains responsible for collision avoidance. The parameters below are the ones the paper leaves symbolic. The social_cost token, filtered and written s, is normalised into ρ, which drives both the field amplitude A and its width σ:

ρ = max(ρfloor, min(1, s / smax))

A = A0 + (Amax − A0) ρ

σ = σ0(1 + 0.58 ρ) · Mgait · Mspeed · kshape

σfront = σ(1 + 0.45 m) for m > m0, else σside = σrear = σ

A floor ρfloor keeps a minimum field even when the token is small, while the hard obstacle core below guarantees geometric clearance independently. Above the floor, the social_cost token, the gait multiplier and the speed multiplier all contribute to the field width.

Amplitude A

  • A0 / Amax: 180 / 252
  • ρfloor: 0.30
  • smax (token cap): 4.0
  • Token gain before filtering: 1.0

Width σ

  • σ0 (base spatial scale): 0.95 m
  • Mgait: quiet 1.32, normal 1.18, urgent 1.06
  • Mspeed: low 1.10, normal 1.00, high 0.92
  • Quiet / low gives the widest field, urgent / high the narrowest

Motion m

  • m0 (moving-person threshold): 0.20
  • Front-lobe coefficient: 0.45
  • Centre offset δ: 0.036 m behind the person
  • Direction taken from observed motion, not estimated body yaw

Group field

  • Agrp = min(Amax, 1.25 A)
  • σlat: 0.65 m
  • σend = max(0.35 m, 0.35 d) for spacing d
  • Merged when pairwise spacing falls in a small range

Hard safety core

  • Outside the social layer, in the obstacle layer
  • Inner safety radius: 0.95 m
  • Inner minimum / maximum cost: 238 / 252
  • Unaffected by the token

Reward

The reward combines velocity tracking with acoustic, gait-style, symmetry and regularization terms. The acoustic terms are token-gated: they scale with the quietness of the gait token, so at a matched speed the quiet token lands softer than the normal or urgent token — from the same network. Falls are handled by episode termination rather than a collision penalty.

Symmetry deserves its own mention. A yaw-rate penalty on straight commands, a left–right joint-mirror term and a contact-force balance term together suppress the lateral drift that a quiet gait otherwise inherits, keeping a straight command straight without giving up the softer footfall.

Group Term What it shapes Weight
TaskLinear velocity trackingExponential on the squared command error1.0
Angular velocity trackingSame form on yaw rate2.5
Move shortfallAnti-stall: progress falling short of the commanded speed−0.9
AcousticToken-scaled foot velocitySquared vertical foot speed, capped−0.55
Phase drop / raiseSoft landing near touchdown, small bonus for lifting1.0
Peak contact-force capContact force above 120 N−1.0
Contact loading rateForce jumping by more than 10 N in one step−2×10−3
Gait styleContact-schedule trackingStance leg loaded, swing leg unloaded, on the token's phase0.8
Foot clearanceSwing height up to 28 mm2.0
Foot liftSecondary clearance term up to 35 mm0.3
Feet air timeAir time beyond 0.5 s at touchdown0.125
Stance widthLateral foot offset up to 0.17 m2.5
SymmetryStraight-yaw penaltyYaw rate while the command is straight−30.0
L/R joint mirrorThigh and calf mismatch between left and right−2.0
L/R force balanceContact-force mismatch between left and right−2×10−4
PostureFore–aft pitchPitch away from flat−8.0
OrientationRoll and pitch of the gravity vector−0.08
Vertical velocityBase bobbing−5.0
Angular velocity (xy)Roll and pitch rate−0.3
SmoothnessAction boundActions driven past their working range−0.1
Action rateChange between consecutive actions−0.02
Joint torqueSquared torque−2×10−4
Joint powerTorque times joint velocity−2×10−5
Joint accelerationSquared joint acceleration−2.5×10−7

Positive weights reward, negative weights penalize. Phase drop / raise is a single signed term whose net effect is a soft-landing penalty.

Training and deployment

Observation (64-d)

  • Angular velocity 3, gravity 3
  • Velocity command 3
  • Token descriptor 7
  • q 12, q̇ 12, previous action 12
  • Gait clock 8, desired contact 4
  • No base linear velocity

Optimization

  • PPO, 4,096 environments × 24 steps
  • γ=0.99, λ=0.95, clip 0.2
  • Adaptive LR, KL target 0.01
  • Entropy 0.01, value coef 0.5
  • ≈25k iterations (≈2.4×109 steps)

Deployment

  • Action scale 0.125 hip, 0.25 thigh/calf
  • Payload 1.0–1.4 kg with CoM shift
  • Friction, gain and push randomization
  • Commands below 0.2 m/s snap to a stand
  • Turning issued at ≥0.4 m/s

BibTeX

@inproceedings{anonymous2027tacet,
  title     = {TACET: Context-Appropriate Acoustic-Social Navigation for Quadrupeds},
  author    = {Anonymous Authors},
  booktitle = {Proc. IEEE Int. Conf. Robotics and Automation (ICRA)},
  note      = {Submitted},
  year      = {2027}
}