A legged robot's locomotion noise is a social variable. TACET infers the social context of the scene
and emits one compact behavior token, 〈gait, speed, social_cost〉,
that sets both where the robot goes and how loudly it walks.
9.3 dBA
quieter than the default gait, at matched speed
100%
personal-space compliance, every scenario
≤2.9 dBA
added at the listener's own position
7
real-world encounters on a Unitree Go2
Abstract
Quadruped robots entering hospitals, care homes, and quiet offices must be context-appropriate not only in
where they move but in how loudly they move: a legged robot's locomotion noise, dominated by
foot–ground impacts, is itself a social variable. Prior social navigation respects human space but treats
the robot as acoustically uniform, while quiet-locomotion methods reduce noise to an operator-specified, context-blind level.
We present TACET, a context-appropriate acoustic-social navigation method that infers social context
from the robot's egocentric view and decides both where it walks and how loudly, coupling a slow fine-tuned
vision-language reasoner to a fast reactive controller through a single compact behavior token,
〈gait, speed, social_cost〉. The same token conditions both a social
costmap (where to go) and a quiet locomotion policy (how loudly to move), while a structured
out-of-view memory keeps recently seen people in the reasoner's context after they leave the camera view.
On a real quadruped, context-conditioned locomotion lowers locomotion noise by up to 9.3 dBA at matched
speed, and across our scenarios the full method keeps personal-space
compliance at 100% with low acoustic intrusion (≤2.9 dBA), jointly improving spatial and acoustic
performance in the evaluated scenarios.
A Dual-Process Architecture
01
Read the social scene
A fine-tuned vision-language model reads the egocentric view and emits one behavior token at roughly
0.25 Hz — slow enough to reason, far outside the safety loop.
02
One token, two mechanisms
The same token scales an asymmetric proxemic costmap and conditions a quiet locomotion policy, so room
and quietness are allocated by a single inferred signal.
03
Remember who left the frame
A structured out-of-view memory carries recently seen people across field-of-view loss, each tagged with
an age-decayed confidence.
The semantic reasoner (person tracker → social memory → VLM reasoner) emits the behavior token;
the reactive controller (social costmap → path planner → velocity controller → locomotion policy)
consumes it to produce socially aware, acoustically controlled motion. The token is the sole interface: the reasoner never outputs
motor commands, and the controller never performs semantic reasoning.
The semantic reasoner
A fine-tuned Qwen2.5-VL-7B maps the egocentric RGB image, a natural-language instruction, and a structured
social memory to a short chain-of-thought followed by the behavior token, with gait ∈
{quiet, normal, urgent}, speed ∈ {low, normal, high}, and social_cost a scalar
amplitude at a resolution of 0.5. Two categorical fields and one scalar make the output directly parseable,
in contrast to free-form language or continuous embeddings.
Supervision comes from auto-labeling: a large VLM produces structured 〈thought, action token〉 labels
on 3,950 egocentric frames from MuSoHu, JRDB-Social, SCAND and InCrowd-VI, every one human-verified, combined with
18,042 COCO images auto-labeled into the same format and 909 SocialNav-SUB frames remapped onto the token
format, for 34,751 training records over 22,901 distinct frames. On a held-out benchmark of 100
human-labeled SCAND frames, exact agreement rises from 30% / 26% / 14% for the base model to
95% / 96% / 69% after QLoRA fine-tuning. A query takes 2.6 s at the median, so tokens are refreshed
at roughly 0.25 Hz and cached — slow reasoning never enters the safety loop.
Out-of-view social memory
A narrow camera field of view means people the robot was just reacting to leave the frame while remaining well
inside its local planning neighborhood. People are detected with YOLOv8n-pose, localized by fusing depth with
the point cloud, and maintained in the odometry frame under a hold / coast / refresh policy: a moving person is
coasted at constant velocity, a stationary person held until a time-to-live expires. Each retained person
becomes a structured record whose confidence starts from source quality and decays with observation age.
Records enter the prompt as reasoning context, never as a rule-based override.
Reactive control — costmap and gait, from the same token
Each person induces an anisotropic Gaussian in their own frame, written into one layer of a layered costmap.
What differs from prior proxemic costmaps is not the shape but where its parameters come from: in TACET the
amplitude and width follow the emitted token, so the same decision that sets how quietly the robot walks also
sets how much room it leaves. Collisions stay in a separate hard obstacle layer, so social cost modulates
preference without compromising geometric safety.
Because quadruped locomotion noise is dominated by foot–ground impacts, whose acoustic energy grows with
foot velocity at contact, TACET reduces perceived noise through control rather than hardware. The
gait and speed tokens enter the policy as physical quantities rather than a
one-hot code: gait selects a contact schedule, a four-beat crawl for quiet and a trot with
decreasing duty factor for normal and urgent, and speed sets its period. One network therefore
covers every gait–speed pair the reasoner emits. A token-gated foot-impact penalty tightens as the gait
token moves from urgent to quiet at matched speed. Unlike an operator-set quiet factor, quietness is
allocated by context rather than preset.
Per-person social fields at the maximum social_cost token: circular (static), egg-shaped
elongated along travel (moving), and one enlarged field containing an interacting pair, protecting the
shared interaction space.
The Gait as a Noise Source
We first characterize the locomotion policy in isolation, with no people present, which separates the gait's
acoustic contribution from that of any detour. Baselines are compared at matched actual speed taken from
LiDAR-inertial odometry, so a method cannot appear quieter merely by moving more slowly.
Speed (m/s)
Method
MNL (dBA) ↓
PNL (dBA) ↓
0.5
Default Go2
68.9 ± 2.8
72.5 ± 0.7
MUTE (β=1)
67.8 ± 2.0
71.5 ± 0.5
Ours
64.1 ± 1.3
65.9 ± 0.8
0.8
Default Go2
74.4 ± 2.4
77.2 ± 1.2
MUTE (β=1)
71.7 ± 2.9
74.6 ± 0.7
Ours
65.1 ± 2.1
68.1 ± 1.6
1.1
Default Go2
74.8 ± 1.9
76.9 ± 1.2
MUTE (β=1)
72.3 ± 2.5
76.6 ± 0.2
Ours
70.6 ± 2.4
75.1 ± 0.6
Speed is the commanded speed. Mean ± standard deviation over ten trials per speed; lower is better.
Ours is the quietest at every matched speed, lowering mean noise level by up to 9.3 dBA over the
default controller and 6.6 dBA over MUTE. Standing still with only its cooling fan running, the Go2 already
registers about 60 dBA at the same microphone position — the floor these levels should be read
against. The margin is smallest at 1.1 m/s, where impact noise dominates and little headroom remains.
Seven Encounters
Nothing in the instruction names a gait, a speed, or an amount of room. The token follows the scene, and the
behavior follows the token — from a quiet crawl past someone working to an urgent trot past a siren light.
With a person in the robot's path
Evaluated with the full set of spatial and acoustic social metrics.
Focused person · workingindoor
Someone absorbed in a stationary task in a quiet room. The robot must reach a nearby goal without acoustically intruding.
quiet · low · 4.0 3.61 m · 1.6 dBA
Focused person · headphonesindoor
The same person, seat, posture and layout. Headphones are the only changed cue — and the gait token changes with them.
normal · low · 4.0 3.62 m · 2.9 dBA
Standing groupindoor
A small group converses. The robot must reach a goal beyond them without cutting through the shared interaction space.
quiet · low · 3.5 2.25 m · 2.7 dBA
Moving personoutdoor
A person walks through the space, where a higher ambient level masks part of the robot's own noise. The robot yields room as it passes.
quiet · normal · 2.5 1.44 m · 1.7 dBA
With no one in the path
The social metrics do not apply here; what matters is that caution is not spent where it buys nothing.
Empty corridorindoor
No one present, so any detour or slowdown is pure overhead.
normal · normal · 1.5 14.4 s to goal
Open sports fieldoutdoor
A large open space with people active far from the robot's path.
normal · high · 1.0 14.0 s to goal
Emergencyindoor
A flashing siren light in a dim corridor. Seeing it, the reasoner decides speed outweighs restraint.
urgent · high · 1.0 11.2 s to goal
Results
PSC is the fraction of the run spent at least 1.2 m from every person; MinDist the closest approach;
ΔL the A-weighted level the robot adds at the listener; AID the fraction of the run for which
ΔL exceeds 5 dBA.
Scenario
Method
PSC (%) ↑
MinDist (m) ↑
ΔL (dBA) ↓
AID (%) ↓
Focused (working)
RPP
95.6
1.03 ± 0.12
8.4 ± 0.5
66.3
SFM
92.7
0.41 ± 0.05
5.5 ± 0.3
51.8
MUTE
97.4
0.91 ± 0.14
6.4 ± 0.3
61.2
VLM-SN
100.0
1.35 ± 0.06
5.9 ± 0.2
58.3
Ours
100.0
3.61 ± 0.17
1.6 ± 0.2
0.0
Focused (headph.)
RPP
95.8
1.06 ± 0.14
8.1 ± 0.6
65.9
SFM
92.2
0.39 ± 0.06
5.7 ± 0.2
52.3
MUTE
97.1
0.89 ± 0.15
6.3 ± 0.3
60.1
VLM-SN
100.0
1.35 ± 0.04
5.7 ± 0.2
57.9
Ours
100.0
3.62 ± 0.19
2.9 ± 0.2
26.5
Standing group
RPP
92.3
0.90 ± 0.19
12.1 ± 1.1
100.0
SFM
92.8
1.02 ± 0.02
7.8 ± 0.6
73.3
MUTE
93.7
1.13 ± 0.05
4.9 ± 0.5
42.9
VLM-SN
90.4
0.87 ± 0.16
7.3 ± 0.5
65.7
Ours
100.0
2.25 ± 0.21
2.7 ± 0.3
22.3
Moving person
RPP
87.7
0.59 ± 0.05
7.9 ± 0.7
83.2
SFM
90.2
0.75 ± 0.04
5.1 ± 0.5
51.3
MUTE
98.9
1.21 ± 0.11
3.8 ± 0.3
24.7
VLM-SN
94.1
0.69 ± 0.07
4.1 ± 0.4
29.4
Ours
100.0
1.44 ± 0.12
1.7 ± 0.3
0.0
Means over ten trials, with standard deviations given for MinDist and ΔL.
VLM-SN abbreviates VLM-Social-Nav.
Ours is the only method that keeps personal-space compliance at 100% in all four scenarios; VLM-Social-Nav
reaches 100% with the focused person but not with the group or the moving person. Ours also keeps a far wider
margin at the closest approach, with MinDist never below 1.44 m against at most 1.35 m for any baseline. All runs
reached the goal without contacting a person, but not always by the robot's own doing. Three baseline conditions
came within arm's reach of a person: SFM in both focused-person conditions, where it repeatedly stalled and
scraped the wall (0.41 and 0.39 m), and RPP with the moving person, which stayed clear only because the person
stepped aside (0.59 m).
What the reasoner adds
The two focused-person conditions hold the scene fixed and change only the headphones. Methods that consume
positions alone behave near-identically in both, as VLM-Social-Nav does (ΔL 5.9 vs. 5.7 dBA).
Ours separates them, keeping the same room (MinDist 3.61 vs. 3.62 m) while relaxing the acoustic effort
where the listener is less sensitive to it (ΔL 1.6 vs. 2.9 dBA, AID 0.0 vs. 26.5%). With the
closest approach unchanged and only the gait token differing, the acoustic difference is consistent with the
gait effect measured on the controlled traverse.
What it costs
Scenario
Baseline mean (s)
Baseline best (s)
Ours (s)
Focused (working)
31.8
21.2
36.8
Focused (headph.)
31.8
22.0
33.8
Standing group
17.2
13.6
25.6
Moving person
16.7
14.8
21.8
Empty corridor
16.2
14.4
14.4
Sports field
18.3
14.4
14.0
Emergency
15.4
14.2
11.2
Mean time to goal over ten trials; the baseline columns summarize RPP, SFM, MUTE and VLM-Social-Nav.
Bold marks the fastest entry in each row.
Where people need space or quiet, TACET spends at most 1.5× the mean baseline time. Where they do not, it
matches the fastest baseline on an empty corridor and is the fastest method outright on the open field and with
the siren light.
Human Study
We ran a blinded video study with 15 observers to complement the objective results. Each observer rated the 20
short clips formed by the four person-facing scenarios and five methods (RPP, SFM, MUTE, VLM-Social-Nav, and
TACET), presented in randomized order and without method labels. After every clip, the observer rated Q1–Q4 on a 0–5 agreement scale. All sessions used the
same headphones and playback level so that relative footstep-noise differences were preserved.
Q1The robot maintained a safe and comfortable distance from the person.
Q2The robot appropriately adjusted its behavior to the social situation and surroundings.
Q3The robot’s footstep noise was appropriate for this quiet indoor setting.
Q4The robot did not unnecessarily disturb the person’s work.
Focused person · headphones
quiet indoor setting
Q1The robot maintained a safe and comfortable distance from the person.
Q2The robot appropriately adjusted its behavior to the social situation and surroundings.
Q3The robot’s footstep noise was appropriate for this quiet indoor setting.
Q4The robot behaved appropriately around a person wearing headphones.
Standing group
indoor conversation
Q1The robot maintained a safe and comfortable distance from the group.
Q2The robot appropriately adjusted its behavior to the social situation and surroundings.
Q3The robot’s footstep noise was appropriate near the group’s conversation.
Q4The robot avoided interrupting the group’s shared conversational space.
Moving person
outdoor crossing
Q1The robot maintained a safe and comfortable distance from the pedestrian.
Q2The robot appropriately adjusted its behavior to the social situation and surroundings.
Q3The robot’s footstep noise was appropriate for this outdoor setting.
Q4The pedestrian could continue without having to alter their path to avoid the robot.
Per-question results of the blinded video study. Q1: safe and comfortable distance; Q2: contextual
consideration; Q3: footstep-noise appropriateness; Q4: scenario-specific etiquette. Bars show mean scores on
the 0–5 agreement scale over paired within-observer ratings; error bars show standard errors. Higher is better.
TACET receives the highest agreement in all sixteen scenario–question cells, averaging 4.54
against 2.23 for MUTE, 1.92 for VLM-Social-Nav, 1.07 for SFM and 0.78 for RPP, and its standard errors are the
smallest (0.14 on average against 0.29 for VLM-Social-Nav), so observers agreed about it more than about any
baseline. The margin narrows only with the moving person, where MUTE reaches 3.33–4.07 — the same
scenario in which MUTE is the strongest baseline on the objective metrics, so its quiet gait alone recovers much
of what observers reward.
Ablation
One component is removed at a time on two representative person-facing scenarios. Separately, the top block of each
scenario keeps the architecture fixed and swaps only the reasoner that emits the token, isolating the
contribution of fine-tuning beyond prompt engineering.
Focused personindoor
Clearance is what the social layer buys; quietness is what the trained policy buys. Removing either one
leaves the other intact. The clip then swaps the reasoner for a prompted base VLM, which keeps the
architecture fixed and loses both.
Moving personoutdoor
People pass in and out of the camera view as the robot walks, so this is where the social memory earns
its place. The reasoner swap that follows barely changes the route here, yet the robot turns audible
above ambient for a third of the run.
Scenario
Variant
PSC (%) ↑
MinDist (m) ↑
ΔL (dBA) ↓
AID (%) ↓
Focused (working)
Qwen2.5-VL-7B
100.0
1.35 ± 0.10
4.5 ± 0.5
31.8
InternVL3.5-8B
100.0
1.33 ± 0.12
4.9 ± 0.4
35.1
Ours (full)
100.0
3.61 ± 0.17
1.6 ± 0.2
0.0
w/o social layer
96.5
1.18 ± 0.11
4.3 ± 0.3
29.3
w/o trained policy
100.0
3.47 ± 0.19
7.4 ± 0.5
63.3
w/o social memory
100.0
3.54 ± 0.24
3.9 ± 0.3
18.7
Moving person
Qwen2.5-VL-7B
100.0
1.38 ± 0.11
3.5 ± 0.5
36.1
InternVL3.5-8B
100.0
1.33 ± 0.12
3.8 ± 0.7
39.3
Ours (full)
100.0
1.44 ± 0.12
1.7 ± 0.3
0.0
w/o social layer
94.7
0.89 ± 0.12
3.8 ± 0.3
37.4
w/o trained policy
100.0
1.39 ± 0.13
4.2 ± 0.5
41.5
w/o social memory
100.0
1.42 ± 0.09
2.6 ± 0.4
19.3
Means over ten trials, with standard deviations given for MinDist and ΔL. Rows above each
rule swap only the reasoner that emits the token; rows below remove one component.
Distance and gait are separable. Removing the trained policy keeps the closest approach within one standard
deviation of the full method, yet degrades the acoustic terms the most of any variant, by 5.8 dBA with the
focused person and 2.5 dBA with the moving one, supporting an acoustic contribution from the gait policy
beyond that of the detour. Removing the social layer costs clearance instead, collapsing MinDist
from 3.61 to 1.18 m and from 1.44 to 0.89 m, both inside the 1.2 m zone the full method never
enters. Removing social memory barely alters the route, yet the robot becomes audible above ambient for 18.7%
and 19.3% of the run: without it, a person who briefly leaves the camera view is no longer carried in the
reasoner's context.
Swapping the reasoner degrades both axes. With the focused person, a prompted Qwen2.5-VL-7B keeps only
1.35 m of clearance against 3.61 m for the full method, close to what removing the social layer
gives — and raises ΔL to 4.5 dBA with AID at 31.8%; InternVL3.5-8B behaves the same way
(1.33 m, 4.9 dBA, 35.1%). With the moving person the route changes far less (1.38 and 1.33 m
against 1.44 m), yet the robot is audible above ambient for 36.1% and 39.3% of the run against 0.0% for the
full method. Since the architecture is unchanged, the loss follows from the token interface itself rather than
from any component.
Implementation
Platform
Unitree Go2 quadruped
Intel RealSense D435i RGB-D
Hesai XT-16 LiDAR
Point-LIO odometry
Nav2 (Smac Hybrid-A*) + pure pursuit
Reasoner
Qwen2.5-VL-7B, 4-bit QLoRA, rank 16
One epoch, learning rate 2×10−5
34,751 records / 22,901 frames
Median query 2.6 s (p95 3.5 s)
Token refresh ≈0.25 Hz, cached
NVIDIA RTX A6000 (48 GB)
Locomotion policy
MLP [512, 256, 128], ELU
50 Hz, joint PD (Kp=40, Kd=1)
Six gait/speed modes, one network
Isaac Lab, deployed through rl_sar
Social costmap
Each tracked person contributes a bounded, anisotropic social field that changes the Nav2 path preference;
the obstacle layer remains responsible for collision avoidance. The parameters below are the ones the paper
leaves symbolic. The social_cost token, filtered and written s, is normalised into
ρ, which drives both the field amplitude A and its width σ:
ρ = max(ρfloor, min(1, s / smax))
A = A0 + (Amax − A0) ρ
σ = σ0(1 + 0.58 ρ) · Mgait · Mspeed · kshape
σfront = σ(1 + 0.45 m) for m > m0, else σside = σrear = σ
A floor ρfloor keeps a minimum field even when the token is small, while the hard obstacle
core below guarantees geometric clearance independently. Above the floor, the social_cost token,
the gait multiplier and the speed multiplier all contribute to the field width.
Amplitude A
A0 / Amax: 180 / 252
ρfloor: 0.30
smax (token cap): 4.0
Token gain before filtering: 1.0
Width σ
σ0 (base spatial scale): 0.95 m
Mgait: quiet 1.32, normal 1.18, urgent 1.06
Mspeed: low 1.10, normal 1.00, high 0.92
Quiet / low gives the widest field, urgent / high the narrowest
Motion m
m0 (moving-person threshold): 0.20
Front-lobe coefficient: 0.45
Centre offset δ: 0.036 m behind the person
Direction taken from observed motion, not estimated body yaw
Group field
Agrp = min(Amax, 1.25 A)
σlat: 0.65 m
σend = max(0.35 m, 0.35 d) for spacing d
Merged when pairwise spacing falls in a small range
Hard safety core
Outside the social layer, in the obstacle layer
Inner safety radius: 0.95 m
Inner minimum / maximum cost: 238 / 252
Unaffected by the token
Reward
The reward combines velocity tracking with acoustic, gait-style, symmetry and regularization
terms. The acoustic terms are token-gated: they scale with the quietness of the
gait token, so at a matched speed the quiet token lands softer than the normal or
urgent token — from the same network. Falls are handled by episode termination rather than
a collision penalty.
Symmetry deserves its own mention. A yaw-rate penalty on straight commands, a left–right
joint-mirror term and a contact-force balance term together suppress the lateral drift that a
quiet gait otherwise inherits, keeping a straight command straight without giving up the
softer footfall.
Group
Term
What it shapes
Weight
Task
Linear velocity tracking
Exponential on the squared command error
1.0
Angular velocity tracking
Same form on yaw rate
2.5
Move shortfall
Anti-stall: progress falling short of the commanded speed
−0.9
Acoustic
Token-scaled foot velocity
Squared vertical foot speed, capped
−0.55
Phase drop / raise
Soft landing near touchdown, small bonus for lifting
1.0
Peak contact-force cap
Contact force above 120 N
−1.0
Contact loading rate
Force jumping by more than 10 N in one step
−2×10−3
Gait style
Contact-schedule tracking
Stance leg loaded, swing leg unloaded, on the token's phase
0.8
Foot clearance
Swing height up to 28 mm
2.0
Foot lift
Secondary clearance term up to 35 mm
0.3
Feet air time
Air time beyond 0.5 s at touchdown
0.125
Stance width
Lateral foot offset up to 0.17 m
2.5
Symmetry
Straight-yaw penalty
Yaw rate while the command is straight
−30.0
L/R joint mirror
Thigh and calf mismatch between left and right
−2.0
L/R force balance
Contact-force mismatch between left and right
−2×10−4
Posture
Fore–aft pitch
Pitch away from flat
−8.0
Orientation
Roll and pitch of the gravity vector
−0.08
Vertical velocity
Base bobbing
−5.0
Angular velocity (xy)
Roll and pitch rate
−0.3
Smoothness
Action bound
Actions driven past their working range
−0.1
Action rate
Change between consecutive actions
−0.02
Joint torque
Squared torque
−2×10−4
Joint power
Torque times joint velocity
−2×10−5
Joint acceleration
Squared joint acceleration
−2.5×10−7
Positive weights reward, negative weights penalize. Phase drop / raise is a single signed term
whose net effect is a soft-landing penalty.
Training and deployment
Observation (64-d)
Angular velocity 3, gravity 3
Velocity command 3
Token descriptor 7
q 12, q̇ 12, previous action 12
Gait clock 8, desired contact 4
No base linear velocity
Optimization
PPO, 4,096 environments × 24 steps
γ=0.99, λ=0.95, clip 0.2
Adaptive LR, KL target 0.01
Entropy 0.01, value coef 0.5
≈25k iterations (≈2.4×109 steps)
Deployment
Action scale 0.125 hip, 0.25 thigh/calf
Payload 1.0–1.4 kg with CoM shift
Friction, gain and push randomization
Commands below 0.2 m/s snap to a stand
Turning issued at ≥0.4 m/s
BibTeX
@inproceedings{anonymous2027tacet,
title = {TACET: Context-Appropriate Acoustic-Social Navigation for Quadrupeds},
author = {Anonymous Authors},
booktitle = {Proc. IEEE Int. Conf. Robotics and Automation (ICRA)},
note = {Submitted},
year = {2027}
}