How to Control an AgileX PiPER Robotic Arm with LLM Model: GPT-6 Astra and Jev

Learn how GPT-6 Astra, Jev, Grounding DINO, SAM 3, and STL geometry were used to control a real AgileX PiPER robotic arm across eight manipulation runs.

Updated on
How to Control an AgileX PiPER Robotic Arm with LLM Model: GPT-6 Astra and Jev
What does it really mean to control a robotic arm with a large language model?

It does not usually mean asking an LLM to generate motor currents or command every joint at every millisecond. A practical system is layered: an AI model helps interpret the task and choose actions, perception software locates objects, a grasp planner decides where the gripper should close, and deterministic robotics software converts the plan into safe physical motion.

The open-source RobotKitAI piper-astra-jev project makes that architecture concrete. It connects GPT-6 Astra and Jev to a real AgileX Robotics PiPER robotic arm and explores eight manipulation runs—from moving a cube into a tray to aligning two reflective metal parts with only about 2 mm of clearance on each side.

For beginners, the project is a useful guide to the building of LLM robot control: multimodal reasoning, object detection, segmentation, depth sensing, grasp planning, known 3D geometry, kinematics, hardware control, and physical feedback.
Important: These are eight single demonstration runs, not a formal benchmark. Each notebook was run once for its recording, so the results should not be treated as a statistical comparison of GPT-6 Astra, Jev, Grounding DINO, or SAM 3. Timings and outcomes can vary between runs.

What Is the PiPER + GPT-6 Astra + Jev Project?

The project's goal is easy to state:
See an object → understand the task → choose an action → plan a grasp → move the arm → verify the result.
The interesting part is that the eight runs do not all solve this problem in the same way.
  • In Run 1, GPT-6 Astra receives camera images directly and acts as the high-level decider.
  • In Runs 2–8, Jev acts as the decider while local perception components provide more structured information about the workspace.
  • Grounding DINO supplies object detections in Runs 2 and 4.
  • SAM 3 supplies segmentation masks in Runs 3 and 5–8.
  • STL geometry is added in Runs 6–8 to reason about balance, recover the pose of reflective objects, and align two parts.
This progression matters more than a simple “GPT-6 vs. Jev” framing. It shows that real-world robot performance depends on the complete manipulation stack—not only on the model selecting the next skill.

Source: X @MarianPogran (https://x.com/MarianPogran/status/2101179679319241076)

 

Hardware and Software Stack

The repository describes the following physical setup.
Component
Role in the project
Why it matters
AgileX PiPER robotic arm
Executes the manipulation skills
Provides the 6-DoF physical platform, gripper, CAN control, and joint feedback
Intel RealSense depth camera
Observes the tabletop and supplies color and depth data
Supports object localization, mask-depth fusion, tray measurements, and visual alignment
CAN interface (can0)
Connects the workstation to PiPER
Carries commands and state between the control software and the arm
Workstation with RTX 4090 or similar
Runs local perception models
Accelerates Grounding DINO and SAM 3 inference
Physical emergency stop
Provides immediate human intervention
Essential because perception or model decisions can be wrong

The software stack is modular rather than monolithic.
Software layer
Component used
Responsibility
High-level decision
GPT-6 Astra or Jev
Interprets the task state and selects an available robot skill
Object detection
Grounding DINO
Returns a bounding box for a text-described target
Object segmentation
SAM 3
Returns the pixels belonging to the target object
Geometry reasoning
STL models and project geometry utilities
Estimates pose from silhouette and searches for geometry-aware grasps
Grasp planning
Box-center, mask-based, or geometry-based search
Chooses where and how the gripper should contact the object
Kinematics
PiPER model from MuJoCo Menagerie
Converts target poses into valid arm configurations
Arm control
piper_control over AgileX piper_sdk
Enables the arm and sends commands to the physical PiPER hardware
Experiment interface
Jupyter notebooks
Guides hardware checks, calibration, dry runs, and live execution

PiPER itself is a lightweight 6-DoF arm with Python API support and compatibility with ROS 1 and ROS 2. Developers can review the current specifications and integration options on the official AgileX PiPER product page.

What Are GPT-6 Astra and Jev?

Although both models help decide what PiPER should do, they are different kinds of AI systems.

What is GPT-6 Astra?

GPT-6 is OpenAI's family of general-purpose reasoning models, and GPT-6 Astra is the flagship model used in Run 1. According to the official OpenAI model documentation, Astra supports text and image input as well as function calling and structured outputs. That combination allows it to inspect camera images, reason about the scene, and select a robot skill through a software interface.
GPT-6 Astra is still not a motor controller. In this project, it contributes visual understanding and high-level reasoning; calibrated robotics software remains responsible for grasping, kinematics, limits, and physical execution.

What is Jev?

Jev is TypeSafe's first System One model. Strictly speaking, it is not a conventional generative LLM. The TypeSafe documentation describes it as a structured decision model: it evaluates text or structured application state and returns typed answers with probabilities instead of generating open-ended prose.
That makes Jev suitable for focused questions such as which allowed skill to run, whether a condition is true, or how a situation should be scored. In Runs 2–8, specialized vision components first turn the camera observation into useful scene information; Jev then acts as the high-level decider. It does not replace Grounding DINO, SAM 3, the geometry pipeline, or the robot controller.

Model
Model type
Input used in this project
Main role in the PiPER pipeline
GPT-6 Astra
General-purpose multimodal reasoning model
Camera images and task context
Understand the visual scene and select robot actions
Jev
Structured System One decision model
State produced with help from specialized perception
Select among defined actions using typed decisions


Why model performance is not an absolute comparison

It would be misleading to treat these runs as proof that one model is universally “better.” GPT-6 Astra and Jev receive different inputs and are embedded in different pipelines. A model that interprets images directly is solving a different problem from a decision model that receives information already extracted by Grounding DINO, SAM 3, depth processing, or STL matching.
Robot outcomes also depend on calibration, lighting, object pose, perception quality, grasp planning, motion constraints, latency, and physical contact. A stronger reasoning model cannot recover information that a sensor never captured, while a specialized decision model may be preferable when an application needs constrained outputs, predictable integration, or lower decision overhead.

The right question is therefore not “Which model wins?” but “Which combination of model, perception, and control components is appropriate for this task?” The eight project runs are demonstrations of different architectures — not that benchmark, but a good example for hands-on learning and experiments.

How Does an LLM Actually Control a Robotic Arm?

The phrase “LLM-controlled robot” can create the wrong mental picture. The model is not the entire controller, and it should not be treated as one.
In this project, the AI model sits at the decision layer. It receives either camera images or processed scene information, reasons about the current task, and chooses from robot skills exposed by the software. The robotics stack then checks and executes the selected skill.

A simplified control loop looks like this:
Task instruction
↓
Camera observation
↓
Perception: direct vision, detection, segmentation, or geometry fitting
↓
Structured scene state
↓
GPT-6 Astra or Jev selects a robot skill
↓
Grasp planning + kinematics + safety checks
↓
PiPER executes the motion over CAN
↓
Camera and gripper feedback update the state
↺

  1. The model decides what to do

At the highest level, the model reasons about task progress. Has the target been found? Should the arm look again, attempt a grasp, lift the object, move to the tray, release it, or recover after a failed attempt?
That is a good fit for a language or multimodal model because the problem is contextual. The correct next action depends on the instruction, what the system currently perceives, and what happened during previous steps.
The model should choose from a constrained skill set rather than invent arbitrary hardware commands. This gives the rest of the system a defined interface to validate.
  1. Perception converts pixels into robot-usable information

A raw image is meaningful to a multimodal model, but a physical robot often needs more precise information:
  • Which object is the target?
  • Where is it in the image?
  • Which pixels belong to it?
  • How far is it from the camera?
  • What is its orientation?
  • Which regions can the gripper contact?
The runs explore three answers. 
GPT-6 Astra can examine the complete camera image and interpret the scene directly;
Grounding DINO locates the target object by drawing a rectangular bounding box around it;
SAM 3 produces a more precise segmentation mask that identifies the pixels belonging to the object.
For known objects, an STL file adds a fourth source: prior knowledge of the object's 3D shape.
  1. A grasp planner decides where the fingers should close

Finding an object is not the same as finding a grasp. A box around a cube may be enough to estimate a center grasp. The same rule can fail on a ring, a trowel, or a tool with a long handle: the center of the box might be empty, lie on the blade, or leave no space for both fingers.
The project therefore progresses from box-center grasps to mask-based grasp search and then to geometry-aware grasps. Each step gives the planner more physically useful information.
  1. Kinematics turn intent into reachable motion

Once the software has a target pose, it still needs to determine how the PiPER arm joints can reach it. Kinematics and robot-specific constraints translate an intended end-effector pose into joint motion.
This is traditional robotics work, and it remains necessary even when an advanced model chooses the action. The LLM does not remove calibration, reachability, collision risk, joint limits, speed limits, or workspace boundaries.
  1. The arm driver executes the command

The project drives PiPER through piper_control, a wrapper around the AgileX piper_sdk. The driver handles the hardware enable sequence and communication cases in which the arm stops responding.
This separation is valuable: high-level model output becomes a request to a controlled robot interface, not an unrestricted message sent directly to the motors.
  1. Physical feedback closes the loop

Robots act in a world where objects can slip, move, disappear behind the gripper, or differ from the estimate. After an action, the system must observe what actually happened.
In this project, vision supports scene updates, while the gripper state confirms whether a grasp succeeded. A lift is allowed only after the fingers stop partway in a way that indicates an object is being held.

The key idea is simple:
The AI model helps decide what should happen next. Specialized perception and robotics software determine where and how it can happen, and physical feedback checks whether it really happened.

Eight Real-World PiPER Robot Manipulation Runs

The eight runs form a practical capability ladder. The scene, perception source, grasp method, or available geometric knowledge changes from one stage to the next.
Run
Decider
Perception and prior knowledge
Task
Main lesson
1
GPT-6 Astra
Astra sees camera images directly
Put a cube into a tray
A multimodal model can participate directly in scene understanding and task decisions
2
Jev
Local Grounding DINO
Put a cube into a tray
Perception and decision-making can be separated into inspectable modules
3
Jev
Local SAM 3
Put a cube into a tray
A segmentation mask gives more object-shape information than a box
4
Jev
Local Grounding DINO
Pick up an object by its handle and place it in the tray
A box-center grasp can be a poor match for an irregular object
5
Jev
Local SAM 3
Repeat the handle task with a mask-planned grasp
Mask and depth data can support a grasp across solid, suitably wide material
6
Jev
SAM 3 + the trowel's STL geometry
Hold the trowel where it balances and carry it level
Known geometry and center of mass can improve grasp stability
7
Jev
SAM 3 + the part's STL geometry
Put a chrome pipe fitting into the tray
A silhouette plus known geometry can help when depth sensing fails on reflective metal
8
Jev
SAM 3 + STL geometry for both parts
Slide the fitting over a standing box spanner
Relative visual alignment and geometry support tighter two-part manipulation
Again, the table summarizes demonstrations, not benchmark trials. It is useful for comparing architectures and information flow, but not for ranking model accuracy, speed, or reliability.

Runs 1–3: From Direct Vision to Specialized Robot Perception

Runs 1–3 use the same cube-to-tray task but change how the workspace is perceived.

Run 1: GPT-6 Astra sees the workspace directly

GPT-6 Astra receives the camera images directly, interprets the cube and tray, and helps select the next robot skill. Calibration, grasp generation, kinematics, limits, and motion control still remain in the robotics stack.

Run 2: Jev receives Grounding DINO detections

Grounding DINO answers “Where is the requested object?” Jev uses that structured result to decide “What should the robot do next?” This separation makes the perception and decision layers easier to inspect, replace, and debug independently.

Run 3: Jev receives a SAM 3 object mask

A bounding box places the object inside a rectangle; a SAM 3 mask identifies its pixels. Either may work for a cube, but the mask preserves more shape information for later, less regular objects.

Runs 4–5: Why Object Detection Is Not Enough for Grasping

Runs 4 and 5 replace the cube with a tool that should be picked up by its handle, showing why localization and grasp planning are different problems.

Run 4: A bounding box leads to a center grasp

Grounding DINO provides a box, so the straightforward grasp is its center. That can work for a cube but fail on a tool, where the center may fall on the blade or another unsuitable point.

Run 5: A segmentation mask supports a physical grasp search

SAM 3 provides a target mask. Fused with depth, it lets the planner search for solid material and enough room for both fingers—moving from “Where is the object?” to “Where can the gripper hold it?”

This does not make one perception model universally better: a box may be enough for localization, while a mask is more informative for contact planning.

Runs 6–8: Adding 3D Geometry for More Precise Manipulation

Runs 6–8 combine the observed silhouette from segmentation with an STL model of the object's expected 3D shape.

Run 6: Grasping a trowel near its balance point

An object grasped away from its center of mass tends to rotate. Using the trowel's geometry and center-of-mass metadata, the planner favors a valid contact line near the ferrule, where the blade-heavy tool balances more naturally and can be carried closer to level.

Run 7: Using geometry when depth fails on chrome

Chrome can produce incorrect or missing depth values. Run 7 matches the fitting's SAM 3 silhouette against rendered views of its STL model. The tabletop fixes its height, and the best silhouette match estimates the remaining pose for a geometry-based grasp and reach.

Run 8: Aligning two reflective parts with about 2 mm to spare

PiPER slides the chrome fitting over a standing box spanner. Its 25.5 mm bore and the spanner's 21.3 mm width leave roughly 2.1 mm per side when centered. Both reflective parts use STL geometry; the camera measures them together and corrects their relative alignment so common calibration error largely cancels. The project reports alignment within 0.8 mm before lowering at 5 mm/s and releasing the fitting.

What Changed from Run 1 to Run 8?

The runs gradually increase the physical knowledge available to the system.

Stage
What the robot stack knows
Question it can answer
Run 1
Camera images and multimodal context
What is in the scene, and what should happen next?
Run 2
Detected target location
Where is the requested object?
Run 3
Segmented target pixels
Which visible pixels belong to it?
Runs 4–5
Candidate contact regions
Where can the gripper close on it?
Run 6
Known geometry and center of mass
Where can the gripper hold it more stably?
Run 7
Silhouette, known geometry, and tabletop constraint
What is the pose when depth is unreliable?
Run 8
Geometry and relative alignment of two objects
How can one part be aligned and placed over another?

The lesson is not that every robotics project needs every component. It is that the required representation depends on the physical task. A cube pick-and-place demo may tolerate a rough center estimate. A reflective-part insertion task needs much tighter geometry, pose, and feedback.

PiPER enabled by GPT-6 Astra and Jev

How Robot Grasping Works: Look, Reach, Touch

The fixed RealSense camera sees the complete workspace only while the arm is out of the way. As the gripper approaches a target, its fingers begin to hide the object. Continuously trusting every new image can then make the system follow part of its own gripper instead of the target.
The project handles this with a three-stage strategy that resembles human reaching.

Stage
What the system does
Trusted information
Look
Observe the target while the view is clear and plan the grasp
Camera perception, depth, mask, and/or geometry
Reach
Move toward the stored grasp even after the fingers occlude the object
Remembered target pose and robot motion
Touch
Check whether the fingers stopped on an object before allowing the lift
Gripper state rather than the image


Look

The robot plans the grasp while the target is fully visible. It keeps that plan when the fingers enter the image. The system also remembers the tray location while the arm is above it.

Reach

During the final approach, a new detection above the table or larger than the object can be a sliver of the target seen on the gripper—or the gripper itself. The system rejects observations that violate configured height or pixel-size limits and trusts the stored plan through the brief blind phase.

Touch

The gripper supplies the success signal. If the fingers close fully, they probably caught nothing. If they stop partway, Arm.holding() indicates that an object is between the jaws, and only then may the arm lift.
This approach does have a limit: if the target moves after it becomes hidden, the remembered plan becomes wrong. The grasp can close on empty space, after which the decider must open the gripper and look again. The project notes that a wrist camera could reduce this blind interval.

How to Run the LLM Project on PiPER Robotic Arm

The repository supplies eight Jupyter notebooks, one for each demonstration. The commands below summarize the published setup; read the repository's current instructions and every notebook's safety checks before enabling real motion.
  1. Prepare the Python environment

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install -e .
  1. Configure credentials locally

Copy .env.example to .env and provide only the credentials needed by the selected run:
OPENROUTER_API_KEY= # Jev runs
OPENAI_API_KEY= # GPT-6 Astra run
HF_TOKEN= # gated SAM 3 model
The project states that .env is ignored by Git. Never commit API keys to a repository or notebook.
  1. Bring up the PiPER CAN interface

sudo ip link set can0 up type can bitrate 1000000
  1. Launch the jupyter notebooks

The notebooks are numbered from the basic Astra cube task through the final fitting-over-spanner task. Each follows the same general sequence: verify the hardware, calibrate the camera, run the safety checks without torque, and only then execute the task on the enabled arm.

For your first experiment, start with Runs 1-3 before moving on to handle grasping, STL-based pose estimation, or the two-part placement task. The later runs introduce additional assumptions and failure modes that are easier to understand once the basic perception-decision-execution loop is working.

Safety: Connecting an AI Model to a Real Robot

A chatbot mistake produces a bad answer. A robot-control mistake can produce unexpected physical motion. The project reduces risk rather than eliminate it. A model may choose the wrong skill, perception may place an object incorrectly, and the arm can move in an unexpected way.

Before every run:
  • Keep a physical emergency stop within immediate reach.
  • Keep hands, clothing, and cables outside the arm's reachable workspace while it is enabled.
  • Remove anything from the table that cannot safely be knocked over.
  • Start at speed_percent=10 and increase speed only after a clean run.
  • Complete the notebook's dry checks before enabling torque.
  • Never leave the enabled arm unattended.
  • Keep fingers out of the gripper jaws; they close with real force.
Also consider task-specific safeguards such as workspace bounds, joint and Cartesian limits, allowlisted skills, speed and force limits, pre-motion validation, timeouts, state monitoring, and a human-controlled enable sequence.

What Robotics Developers Can Learn from the Project

  1. LLM is the part of a hybrid system

The model is most useful as one component in a hybrid system. Specialized perception, deterministic motion code, safety checks, and physical feedback remain essential.
  1. The best perception output depends on the action

A bounding box can be enough to select or roughly locate an object. A segmentation mask is more useful when the grasp depends on visible shape. A known 3D model becomes valuable when balance, hidden surfaces, or reflective materials matter.
  1. Robot failures should be diagnosable by layer

Separating perception, decision, grasp planning, kinematics, and execution makes debugging more practical. “The robot missed” is not a diagnosis. The useful question is whether it saw the wrong target, estimated the wrong pose, chose a poor grasp, requested an unreachable pose, or lost the object during execution.
  1. Relative measurements can beat absolute calibration

Run 8 uses a powerful robotics pattern. When two objects appear in the same camera frame, their relative offset can be more accurate than two separately transformed world coordinates because common calibration error can cancel.
  1. Memory is part of manipulation

The Look–Reach–Touch strategy shows that a robot cannot assume uninterrupted visibility. It needs short-term state: the last trusted target pose, the planned grasp, the tray location, and the result of the most recent action.
  1. Demonstrations and benchmarks answer different questions

These single runs show that the architectures can be assembled and exercised on real hardware. They do not establish success rates, robustness across lighting and object layouts, latency distributions, or superiority of one model.

A formal evaluation would require repeated trials, controlled variations, defined success criteria, failure categories, timing methodology, and confidence intervals. That would be a valuable next step, but it is outside the claim made by this project.

Why PiPER Robotic Arm Is a Useful Platform for Robot Control Experiments

PiPER robotic arm supports robotics development, exploring robot manipulation and control

PiPER 6-DoF robotic arm brings AI experiments out of simulation and into a manageable physical workspace. Its open development options let researchers combine conventional robot programming with newer perception and reasoning systems.

PiPER is a 6-DoF robotic arm with a 1.5kg payload, 4.2kg body weight, approximately 626mm reach, Python API control, and ROS 1/ROS 2 support. Those characteristics make it relevant to desktop manipulation, education, teleoperation, data collection, and embodied-AI prototyping.

Most importantly, a physical arm exposes problems that a text-only or simulation-only demo can hide:
  • camera-to-robot calibration;
  • self-occlusion;
  • noisy or missing depth;
  • grasp uncertainty;
  • object balance and contact;
  • reach and orientation constraints;
  • hardware latency and communication failures; and
  • real safety boundaries.
Those are not distractions from embodied AI. They are the work embodied AI systems must eventually handle.

Frequently Asked Questions

Can an LLM model control a robotic arm?

Yes, but usually as part of a layered system. An LLM or multimodal model can interpret instructions, reason about a scene, and select robot skills. Perception, grasp planning, kinematics, safety checks, and the hardware driver should still handle the specialized work required to execute those skills.

Did GPT-6 Astra directly control PiPER's motors?

No. In the RobotKitAI project, GPT-6 Astra acts at the high-level decision and visual-reasoning layer. Robot-specific software handles grasp generation, kinematics, motion, safety limits, and communication with PiPER.

What is the difference between GPT-6 Astra and Jev in these runs?

Run 1 gives GPT-6 Astra the camera images directly. Runs 2–8 use Jev as the decider with perception outputs from Grounding DINO or SAM 3, and later add STL geometry. Because each setup was recorded once, the project is not a benchmark comparing the models' reliability or performance.

Why use SAM 3 instead of only Grounding DINO for robot grasping?

Grounding DINO returns a bounding box, which is useful for locating an object. SAM 3 returns an object mask, which can describe the visible shape more precisely. A grasp planner can combine that mask with depth to look for solid material and room for both gripper fingers. Which method is appropriate depends on the object and task.

Why does the project use STL files?

An STL file supplies known 3D geometry. The project uses geometry to search for a balanced trowel grasp, estimate the pose of reflective chrome parts from their silhouettes, and align two known parts in the final run.

Can a depth camera see chrome objects reliably?

Often not. Reflective surfaces can return incorrect or missing depth. In Run 7, the project estimates the chrome fitting's pose by matching its segmented silhouette to rendered views of its known STL geometry while using the tabletop as a height constraint.

What is Look–Reach–Touch in robot grasping?

It is the project's strategy for handling self-occlusion. The robot looks and plans while the object is visible, reaches using the remembered plan when the gripper blocks the camera, and confirms the grasp using gripper feedback before lifting.

Is this project a benchmark of GPT-6 Astra versus Jev?

No. The repository explicitly describes the experiments as single runs rather than a benchmark. They demonstrate different system designs but do not provide enough repeated trials for a statistical model comparison.

What hardware is needed to reproduce the project?

The published setup uses an AgileX PiPER arm connected over CAN, an Intel RealSense depth camera, an RTX 4090-class workstation for local perception, and a physical emergency stop. The exact requirements can vary with the models and runs you choose.

Does AgileX PiPER support ROS 2 and Python development?

Yes. AgileX Robotics lists Python API control and both ROS 1 and ROS 2 support for PiPER. The RobotKitAI project itself uses notebook-driven Python code and CAN communication rather than requiring every component to run through ROS.

From Language Models to Physical Intelligence

The most useful question is no longer simply, “Can an LLM control a robot?”
Of course it can, but the better question is:
Where should a general-purpose model sit inside the robot-control stack, and which responsibilities should remain with specialized perception, geometry, control, and safety systems?
The piper-astra-jev demonstrations offer eight concrete answers. Direct multimodal vision can support a simple task. Structured detections make the pipeline modular. Segmentation improves the information available to a grasp planner. Known geometry helps with balance and reflective surfaces. Relative visual feedback supports tighter placement. Throughout the process, kinematics, hardware interfaces, constraints, and touch feedback turn a high-level decision into a real physical result.

References: Explore LLM and Build with PiPER Robotic Arm

  • Study or reproduce the eight runs: Visit the RobotKitAI piper-astra-jev repository on GitHub for the notebooks, source code, setup details, run documentation, and safety notes. Huge thanks to Michal Kubenka and Marian Pogran for creating this project!
  • Build your own AI robotic arm project: Explore the AgileX Robotics PiPER for specifications, Python/ROS development options, and contact information.
Whether you start with a cube pick-and-place task or a geometry-aware manipulation pipeline, begin slowly, validate each layer independently, and keep a human in control of the physical system.
Updated on