buster/packages/ai/tests/utils/memory/message-history.test.ts

885 lines
29 KiB
TypeScript
Raw Normal View History

Mastra braintrust (#391) * type fixes * biome clean on ai * add user to flag chat * attempt to get vercel deployed * Update tsup.config.ts * Update pnpm-lock.yaml * Add @buster/server2 Hono API app with Vercel deployment configuration * slack oauth integration * mainly some clean up and biome formatting * slack oauth * slack migration + snapshot * remove unused files * finalized docker image for porter * Create porter_app_buster-server_3155.yml file * Add integration tests for Slack handler and refactor Slack OAuth service - Introduced integration tests for the Slack handler, covering OAuth initiation, callback handling, and integration status retrieval. - Refactored Slack OAuth service to improve error handling and ensure proper integration state management. - Updated token storage implementation to use a database vault instead of Supabase. - Enhanced existing tests for better coverage and reliability, including cleanup of test data. - Added new utility functions for managing vault secrets in the database. * docker image update * new prompts * individual tests and a schema fix * server build * final working dockerfile * Update Dockerfile * new messages to slack messages (#369) * Update dockerfile * Update validate-env.js * update build pipeline * Update the dockerfile flow * finalize logging for pino * stable base * Update cors middleware logger * Update cors.ts * update docker to be more imformative * Update index.ts * Update auth.ts * Update cors.ts * Update cors.ts * Update logger.ts * remove logs * more cors updates * build server shared * Refactor PostgreSQL credentials handling and remove unused memory storage. Update package dependencies. (#370) * tons of file parsing errors (#371) * Refactor PostgreSQL credentials handling and remove unused memory storage. Update package dependencies. * tons of file parsing errors * Dev mode updates * more stable electric handler * Dal/agent-self-healing-fixes (#372) * change to 6 min * optmizations around saving and non-blocking actions. * stream optimizations * Dal/agent-self-healing-fixes (#373) * change to 6 min * optmizations around saving and non-blocking actions. * stream optimizations * change porter staging deploy to mastra-braintrust. * new path for porter deploy * deploy to staging fix * Create porter_app_mastra-braintrust-api_3155.yml file (#375) Co-authored-by: porter-deployment-app[bot] <87230664+porter-deployment-app[bot]@users.noreply.github.com> * Update sizing and opacity * supe up the instance for mastra * environment staging * ssl script * copy path * Update list padding * no throttle and the anthropic cached * move select to the top * Update margin inline start * shrink reasoning vertical space to 2px * semi bold font for headers * update animation timing * haiku * Add createTodoList tool and integrate into create-todos-step * chat helper on post chat * only trigger cicd when change made * Start created streaming text components * Refactor analyst agent task to initialize Braintrust logging asynchronously and parallelize database queries for improved performance. Adjusted cleanup timeout for Braintrust traces to reduce delays. * fixed reasoned for X, so that it rounds down to the minute * Update users page * update build pipeline for new web * document title update * Named chats for page * Datasets titles * Refactor visualization tools and enhance error handling in retryable agent stream. Removed unused metricValueLabel from metrics file tool, updated metric configuration schemas, and improved healing mechanism for tool errors during streaming. * analyst * document title updates * Update useDocumentTitle.tsx * Refactor tool choice configuration in create-todos-step to use structured object. Remove exponential backoff logic from retryable agent stream for healable errors. Introduce new test for real-world healing scenarios in retryable agent stream. * Refactor SQL validation logic in modify-metrics-file-tool to skip unnecessary checks when SQL has not changed. Enhance error handling and update validation messages. Clean up code formatting for improved readability. * update collapse for filecard * chevron collapse * Jacob prompt changes (#376) * prompt changes to improve filtering logic and handle priv/sec errors * prompt changes to make aggregation better and improved filter best practices * Update packages/ai/src/steps/create-todos-step.ts Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * Update packages/ai/src/agents/think-and-prep-agent/think-and-prep-instructions.ts Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * Update packages/ai/src/steps/create-todos-step.ts Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> --------- Co-authored-by: Jacob Anderson <jacobanderson@Jacobs-MacBook-Air.local> Co-authored-by: dal <dallin@buster.so> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * think and prep * change header and strong fonts weights * Update get collection * combo chart x axis update * Create a chart schemas as types * schema types * simple unit tests for line chart props * fix the response file ordering iwth active selection. * copy around reasoning messages taken care of * fix nullable user message and file processing and such. * update ticks for chart config * fix todo parsing. * app markdown update * Update splitter to use border instead of width * change ml * If no file is found we should auto redirect * Refactor database connection handling to support SSL modes. Introduced functions to extract SSL parameters and manage connections based on SSL requirements, including a custom verifier for unverified connections. * black box message update * chat title updates * optimizations for trigger. * some keepalive logic on the anthropic cached * keep title empty until new one * no duplicate messages * null user message on asset pull * posthog error handling * 20 sec idle timeout on anthropic * null req message * fixed modificiation names missing * Refactor tool call handling to support new content array format in asset messages and context loaders * cache most recent file from workflow * Enhance date and number detection in createDataMetadata function to improve data type handling for metrics files * group hover effect for message * logging for chat * Add messageId handling and file association tracking in dashboard and metrics tools - Updated runtime context to include messageId in create and modify dashboard and metrics file tools. - Implemented file association tracking based on messageId in create and modify functions for both dashboards and metrics. - Ensured type consistency by using AnalystRuntimeContext in runtime context parameters. * logging for chat * message type update * Route to first file instead * trigger moved to catalog * Enhance file selection logic to support YAML parsing and improve logging - Updated `extractMetricIdsFromDashboard` to first attempt JSON parsing, falling back to a regex-based YAML parsing for metric IDs. - Added detailed debug logging in `selectFilesForResponse` to track file selection process, including metrics and dashboards involved. - Introduced tests for various scenarios in `file-selection.test.ts` to ensure correct behavior with dashboard context and edge cases. * trigger dev v4-beta * Retry + Self Healing (#381) * Refactor retry logic in analyst and think-and-prep steps Co-authored-by: dallin <dallin@buster.so> * some fixes * console log error * self healing * todos retry --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> * remove lots of logs * Remove chat streaming * Remove chat streaming * timeout * Change to updated at field * link to home * Update timeout settings for HTTP and HTTPS agents from 20 seconds to 10 seconds for improved responsiveness. * Add utils module and integrate message conversion in post_chat_handler * Implement error handling for extract values (#382) * Remove chat streaming * Improve error handling and logging in extract values and chat title steps Co-authored-by: dallin <dallin@buster.so> --------- Co-authored-by: Nate Kelley <nate@buster.so> Co-authored-by: Cursor Agent <cursoragent@cursor.com> * loading icon for buster avatar * finalize tooltip cache * upgrade mastra * increase retries * Add redo functionality for chat messages - Introduced `redoFromMessageId` parameter in `handleExistingChat` to allow users to specify a message to redo from. - Implemented validation to ensure the specified message belongs to the current chat. - Added `softDeleteMessagesFromPoint` function to soft delete a message and all subsequent messages in the same chat, facilitating the redo feature. * fix electric potential memory leak * tooltip cache and chart cleanup * Update bullet to be more indented * latest version number * add support endpoint to new server * Fix jank in combo bar charts * index check for dashboard * Collapse only if there are metrics * Is finished reasoing back * Update dependencies and enhance chat message handling - Upgraded `@mastra/core` to version 0.10.8 and added `node-sql-parser` at version 5.3.10 in the lock file. - Improved integration tests for chat message redo functionality, ensuring correct behavior when deriving `chat_id` from `message_id`. - Enhanced error handling and validation in the `initializeChat` function to manage cases where `chat_id` is not provided. * Update pnpm-lock and enhance chat message integration tests - Added `node-sql-parser` version 5.3.10 to dependencies and updated the lock file. - Improved integration tests for chat message redo functionality, ensuring accurate deletion and retrieval of messages. - Enhanced the `initializeChat` function to derive `chat_id` from `message_id` when not provided, improving error handling and validation. * remove .env import breaking build * add updated at to the get chat handler * zmall runtime error fix * permission tests passing * return updated at on the get chat handler now * slq parser fixes * Implement chat access control logic and add comprehensive tests - Developed the `canUserAccessChat` function to determine user access to chats based on direct permissions, collection permissions, creator status, and organizational roles. - Introduced helper functions for checking permissions and retrieving chat information. - Added integration tests to validate access control logic, covering various scenarios including direct permissions, collection permissions, and user roles. - Created unit tests to ensure the correctness of the access control function with mocked database interactions. - Included simple integration tests to verify functionality with existing database data. * sql parser and int tests working. * fix test and lint issues * comment to kick off deployment lo * access controls on datasets * electric context bug fix with sql helpers. * permission and read only * Add lru-cache dependency and export cache management functions - Added `lru-cache` as a dependency in the access-controls package. - Exported new cache management functions from `chats-cached` module, including `canUserAccessChatCached`, `getCacheStats`, `resetCacheStats`, `clearCache`, `invalidateAccess`, `invalidateUserAccess`, and `invalidateChatAccess`. * packages deploy as well * wrong workflow lol * Update AppVerticalCodeSplitter.tsx * Add error handling for query run and SQL save operations Co-authored-by: natemkelley <natemkelley@gmail.com> * Trim whitespace from input values before sending chat prompts Co-authored-by: natemkelley <natemkelley@gmail.com> * type in think-and-prep * use the cached access chat * update package version * new asset import message * Error fallback for login * Update BusterChart.BarChart.stories.tsx * Staging changes to fix number card titles, combo chart axis, and using dynamic filters (#386) Co-authored-by: Jacob Anderson <jacobanderson@Jacobs-MacBook-Air.local> * db init command pass through * combo chart fixes (#387) Co-authored-by: Jacob Anderson <jacobanderson@Jacobs-MacBook-Air.local> * clarifying question and connection logic * pino pretty error fix * clarifying is a finishing tool * change update latest version logic * Update support endpoint * fixes for horizontal bar charts and added the combo chart logic to update metrics (#388) Co-authored-by: Jacob Anderson <jacobanderson@Jacobs-MacBook-Air.local> * permission fix on dashboard metric handlers for workspace and data admin * Add more try catches * Hide avatar is no more * Horizontal bar fixes (#389) * fixes for horizontal bar charts and added the combo chart logic to update metrics * hopefully fixed horizontal bar charts --------- Co-authored-by: Jacob Anderson <jacobanderson@Jacobs-MacBook-Air.local> * reasoning shimmer update * Make the embed flow work with versions * new account warning update * Move support modal * compact number for pie label * Add final reasoning message tracking and workflow start time to chunk processor and related steps - Introduced `finalReasoningMessage` to schemas in `analyst-step`, `mark-message-complete-step`, and `create-todos-step`. - Updated `ChunkProcessor` to calculate and store the final reasoning message based on workflow duration. - Enhanced various steps to utilize the new `workflowStartTime` for better tracking of execution duration. - Improved database update logic to include `finalReasoningMessage` when applicable. * 9 digit cutoff for pie * trigger update * test on mastra braintrust * test deployment * testing * pnpm install * pnpm * node 22 * pnpm version * trigger main * get initial chat file * hono main deploymenbt * clear timeouts * Remove console logs * migration test to staging * db url * try again * k get rid of tls var * hmmm lets try this * mark migrations * fix migration file? * drizzle-kit upgrade * tweaks to the github actions --------- Co-authored-by: Nate Kelley <nate@buster.so> Co-authored-by: porter-deployment-app[bot] <87230664+porter-deployment-app[bot]@users.noreply.github.com> Co-authored-by: Nate Kelley <133379588+nate-kelley-buster@users.noreply.github.com> Co-authored-by: Jacob Anderson <jacobanderson@Jacobs-MacBook-Air.local> Co-authored-by: jacob-buster <jacob@buster.so> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: natemkelley <natemkelley@gmail.com>
2025-07-03 05:33:40 +08:00
import type { CoreMessage } from 'ai';
import { describe, expect, test } from 'vitest';
import {
extractMessageHistory,
getAllToolsUsed,
getConversationSummary,
getLastToolUsed,
isToolCallOnlyMessage,
properlyInterleaveMessages,
unbundleMessages,
} from '../../../src/utils/memory/message-history';
import { hasToolCallId, validateArrayAccess } from '../../../src/utils/validation-helpers';
describe('Message History Utilities', () => {
describe('Message Format Validation', () => {
test('should handle properly formatted unbundled messages', () => {
const properMessages: CoreMessage[] = [
{
role: 'user',
content: 'What are our top 5 products by revenue in the last quarter?',
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'sequentialThinking',
toolCallId: 'toolu_01Um5qidhhwormgMx9mASBv2',
args: {
thought: 'I need to analyze the request...',
isRevision: false,
thoughtNumber: 1,
totalThoughts: 1,
needsMoreThoughts: false,
nextThoughtNeeded: false,
},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: { success: true },
toolName: 'sequentialThinking',
toolCallId: 'toolu_01Um5qidhhwormgMx9mASBv2',
},
],
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'submitThoughts',
toolCallId: 'toolu_01LA17JT7CUATsE4YX2Dy8oz',
args: {},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: {},
toolName: 'submitThoughts',
toolCallId: 'toolu_01LA17JT7CUATsE4YX2Dy8oz',
},
],
},
];
const processed = extractMessageHistory(properMessages);
expect(processed).toHaveLength(5);
expect(validateArrayAccess(processed, 0, 'processed messages')?.role).toBe('user');
expect(validateArrayAccess(processed, 1, 'processed messages')?.role).toBe('assistant');
expect(validateArrayAccess(processed, 2, 'processed messages')?.role).toBe('tool');
expect(validateArrayAccess(processed, 3, 'processed messages')?.role).toBe('assistant');
expect(validateArrayAccess(processed, 4, 'processed messages')?.role).toBe('tool');
});
test('should unbundle messages that have mixed content', () => {
const bundledMessages: CoreMessage[] = [
{
role: 'user',
content: 'What are our top 5 products?',
},
{
role: 'assistant',
content: [
{ type: 'text', text: 'Let me analyze that for you.' },
{
type: 'tool-call',
toolName: 'analyzeData',
toolCallId: 'tool-1',
args: { query: 'top 5 products' },
},
{
type: 'tool-call',
toolName: 'generateChart',
toolCallId: 'tool-2',
args: { type: 'bar' },
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: { data: 'product data' },
toolName: 'analyzeData',
toolCallId: 'tool-1',
},
],
},
];
const unbundled = unbundleMessages(bundledMessages);
// Should have: user, assistant (text), assistant (tool-1), assistant (tool-2), tool result
expect(unbundled).toHaveLength(5);
expect(validateArrayAccess(unbundled, 0, 'unbundled messages')?.role).toBe('user');
expect(validateArrayAccess(unbundled, 1, 'unbundled messages')?.role).toBe('assistant');
expect(validateArrayAccess(unbundled, 1, 'unbundled messages')?.content).toEqual([
{ type: 'text', text: 'Let me analyze that for you.' },
]);
expect(validateArrayAccess(unbundled, 2, 'unbundled messages')?.role).toBe('assistant');
const unbundled2 = validateArrayAccess(unbundled, 2, 'unbundled messages');
expect(unbundled2 ? isToolCallOnlyMessage(unbundled2) : false).toBe(true);
expect(validateArrayAccess(unbundled, 3, 'unbundled messages')?.role).toBe('assistant');
const unbundled3 = validateArrayAccess(unbundled, 3, 'unbundled messages');
expect(unbundled3 ? isToolCallOnlyMessage(unbundled3) : false).toBe(true);
expect(validateArrayAccess(unbundled, 4, 'unbundled messages')?.role).toBe('tool');
});
});
describe('Tool Detection', () => {
test('should correctly identify the last tool used', () => {
const messages: CoreMessage[] = [
{
role: 'user',
content: 'Analyze this data',
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'sequentialThinking',
toolCallId: 'tool-1',
args: {},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: {},
toolName: 'sequentialThinking',
toolCallId: 'tool-1',
},
],
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'submitThoughts',
toolCallId: 'tool-2',
args: {},
},
],
},
{
role: 'tool',
content: [
{ type: 'tool-result', result: {}, toolName: 'submitThoughts', toolCallId: 'tool-2' },
],
},
];
const lastTool = getLastToolUsed(messages);
expect(lastTool).toBe('submitThoughts');
});
test('should get all tools used in conversation', () => {
const messages: CoreMessage[] = [
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'searchDataCatalog',
toolCallId: 'tool-1',
args: {},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: {},
toolName: 'searchDataCatalog',
toolCallId: 'tool-1',
},
],
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'analyzeData',
toolCallId: 'tool-2',
args: {},
},
],
},
{
role: 'tool',
content: [
{ type: 'tool-result', result: {}, toolName: 'analyzeData', toolCallId: 'tool-2' },
],
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'searchDataCatalog',
toolCallId: 'tool-3',
args: {},
},
],
},
];
const tools = getAllToolsUsed(messages);
expect(tools).toContain('searchDataCatalog');
expect(tools).toContain('analyzeData');
expect(tools).toHaveLength(2); // Should not duplicate
});
});
describe('Conversation Summary', () => {
test('should correctly summarize conversation with proper message structure', () => {
const messages: CoreMessage[] = [
{
role: 'user',
content: 'First question',
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'think',
toolCallId: 'tool-1',
args: {},
},
],
},
{
role: 'tool',
content: [{ type: 'tool-result', result: {}, toolName: 'think', toolCallId: 'tool-1' }],
},
{
role: 'assistant',
content: 'Here is my response',
},
{
role: 'user',
content: 'Follow up question',
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'analyze',
toolCallId: 'tool-2',
args: {},
},
],
},
{
role: 'tool',
content: [{ type: 'tool-result', result: {}, toolName: 'analyze', toolCallId: 'tool-2' }],
},
];
const summary = getConversationSummary(messages);
expect(summary.userMessages).toBe(2);
expect(summary.assistantMessages).toBe(3); // 2 with tool calls, 1 with text
expect(summary.toolCalls).toBe(2);
expect(summary.toolResults).toBe(2);
expect(summary.toolsUsed).toEqual(['think', 'analyze']);
});
});
describe('Tool Call Only Messages', () => {
test('should identify messages that only contain tool calls', () => {
const toolCallOnlyMessage: CoreMessage = {
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'submitThoughts',
toolCallId: 'tool-1',
args: {},
},
],
};
const mixedMessage: CoreMessage = {
role: 'assistant',
content: [
{ type: 'text', text: 'Here is some text' },
{
type: 'tool-call',
toolName: 'submitThoughts',
toolCallId: 'tool-1',
args: {},
},
],
};
const textOnlyMessage: CoreMessage = {
role: 'assistant',
content: 'Just text content',
};
expect(isToolCallOnlyMessage(toolCallOnlyMessage)).toBe(true);
expect(isToolCallOnlyMessage(mixedMessage)).toBe(false);
expect(isToolCallOnlyMessage(textOnlyMessage)).toBe(false);
});
});
describe('Sequential Message Order Preservation', () => {
test('should preserve exact sequential order from database example', () => {
// This is the exact structure from the user's database - already properly formatted
const databaseMessages: CoreMessage[] = [
{
role: 'user',
content: 'Who is my top customer?',
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolCallId: 'toolu_01LmHSAwa8MeggWntV8gE1fG',
toolName: 'sequentialThinking',
args: {
thought:
'I need to address the TODO list items for this user request about finding their top customer...',
isRevision: false,
thoughtNumber: 1,
totalThoughts: 2,
needsMoreThoughts: false,
nextThoughtNeeded: true,
},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: { success: true },
toolName: 'sequentialThinking',
toolCallId: 'toolu_01LmHSAwa8MeggWntV8gE1fG',
},
],
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolCallId: 'toolu_015T6fk9RhcJ9AuCYDtdsQba',
toolName: 'sequentialThinking',
args: {
thought: 'Since I have no database documentation provided...',
isRevision: false,
thoughtNumber: 2,
totalThoughts: 3,
needsMoreThoughts: false,
nextThoughtNeeded: true,
},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: { success: true },
toolName: 'sequentialThinking',
toolCallId: 'toolu_015T6fk9RhcJ9AuCYDtdsQba',
},
],
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolCallId: 'toolu_01QtPVf5tYPydXeXWGoCKbpH',
toolName: 'executeSql',
args: {
statements: [
"SELECT table_name FROM information_schema.tables WHERE table_schema = 'public' LIMIT 25",
"SELECT column_name, data_type FROM information_schema.columns WHERE table_name LIKE '%customer%' LIMIT 25",
"SELECT column_name, data_type FROM information_schema.columns WHERE table_name LIKE '%order%' LIMIT 25",
],
},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: {
results: [
/* ... */
],
},
toolName: 'executeSql',
toolCallId: 'toolu_01QtPVf5tYPydXeXWGoCKbpH',
},
],
},
];
// Extract message history should NOT modify the structure
const extracted = extractMessageHistory(databaseMessages);
// Should be exactly the same
expect(extracted).toEqual(databaseMessages);
expect(extracted).toHaveLength(7);
// Verify the sequential pattern is preserved
expect(validateArrayAccess(extracted, 0, 'extracted')?.role).toBe('user');
expect(validateArrayAccess(extracted, 1, 'extracted')?.role).toBe('assistant');
expect(validateArrayAccess(extracted, 2, 'extracted')?.role).toBe('tool');
expect(validateArrayAccess(extracted, 3, 'extracted')?.role).toBe('assistant');
expect(validateArrayAccess(extracted, 4, 'extracted')?.role).toBe('tool');
expect(validateArrayAccess(extracted, 5, 'extracted')?.role).toBe('assistant');
expect(validateArrayAccess(extracted, 6, 'extracted')?.role).toBe('tool');
// Verify tool calls and results are properly paired
const toolCallIds = [
'toolu_01LmHSAwa8MeggWntV8gE1fG',
'toolu_015T6fk9RhcJ9AuCYDtdsQba',
'toolu_01QtPVf5tYPydXeXWGoCKbpH',
];
for (let i = 0; i < toolCallIds.length; i++) {
const assistantIdx = 1 + i * 2;
const toolIdx = 2 + i * 2;
// Get tool call ID from assistant message
const assistantMsg = validateArrayAccess(extracted, assistantIdx, 'assistant message');
const assistantContent = assistantMsg?.content;
if (Array.isArray(assistantContent) && assistantContent.length > 0) {
const toolCall = validateArrayAccess(assistantContent, 0, 'tool call');
if (hasToolCallId(toolCall)) {
expect(toolCall.toolCallId).toBe(validateArrayAccess(toolCallIds, i, 'tool call id'));
}
}
// Verify matching tool result
const toolMsg = validateArrayAccess(extracted, toolIdx, 'tool message');
const toolContent = toolMsg?.content;
if (Array.isArray(toolContent) && toolContent.length > 0) {
const toolResult = validateArrayAccess(toolContent, 0, 'tool result');
if (hasToolCallId(toolResult)) {
expect(toolResult.toolCallId).toBe(validateArrayAccess(toolCallIds, i, 'tool call id'));
}
}
}
});
test('should handle messages bundled incorrectly (the bug scenario)', () => {
// This represents what might come from the AI SDK if it bundles messages
const bundledMessages: CoreMessage[] = [
{
role: 'user',
content: 'Who is my top customer?',
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolCallId: 'toolu_1',
toolName: 'think',
args: { thought: 'First thought' },
},
{
type: 'tool-call',
toolCallId: 'toolu_2',
toolName: 'analyze',
args: { data: 'customers' },
},
{
type: 'tool-call',
toolCallId: 'toolu_3',
toolName: 'finalize',
args: { result: 'done' },
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
toolCallId: 'toolu_1',
toolName: 'think',
result: { success: true },
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
toolCallId: 'toolu_2',
toolName: 'analyze',
result: { success: true },
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
toolCallId: 'toolu_3',
toolName: 'finalize',
result: { success: true },
},
],
},
];
// extractMessageHistory should now fix the bundling
const extracted = extractMessageHistory(bundledMessages);
// Should have been properly interleaved
expect(extracted).toHaveLength(7); // user + 3*(assistant + tool)
// Verify the sequential pattern
expect(validateArrayAccess(extracted, 0, 'extracted')?.role).toBe('user');
expect(validateArrayAccess(extracted, 1, 'extracted')?.role).toBe('assistant');
expect(validateArrayAccess(extracted, 2, 'extracted')?.role).toBe('tool');
expect(validateArrayAccess(extracted, 3, 'extracted')?.role).toBe('assistant');
expect(validateArrayAccess(extracted, 4, 'extracted')?.role).toBe('tool');
expect(validateArrayAccess(extracted, 5, 'extracted')?.role).toBe('assistant');
expect(validateArrayAccess(extracted, 6, 'extracted')?.role).toBe('tool');
// Verify each assistant message has only one tool call
const extracted1 = validateArrayAccess(extracted, 1, 'extracted');
expect(extracted1?.content).toHaveLength(1);
const content1 = extracted1?.content;
if (
Array.isArray(content1) &&
content1[0] &&
typeof content1[0] === 'object' &&
'toolCallId' in content1[0]
) {
expect(content1[0].toolCallId).toBe('toolu_1');
}
const extracted3 = validateArrayAccess(extracted, 3, 'extracted');
expect(extracted3?.content).toHaveLength(1);
const content3 = extracted3?.content;
if (
Array.isArray(content3) &&
content3[0] &&
typeof content3[0] === 'object' &&
'toolCallId' in content3[0]
) {
expect(content3[0].toolCallId).toBe('toolu_2');
}
const extracted5 = validateArrayAccess(extracted, 5, 'extracted');
expect(extracted5?.content).toHaveLength(1);
const content5 = extracted5?.content;
if (
Array.isArray(content5) &&
content5[0] &&
typeof content5[0] === 'object' &&
'toolCallId' in content5[0]
) {
expect(content5[0].toolCallId).toBe('toolu_3');
}
});
});
describe('properlyInterleaveMessages', () => {
test('should interleave bundled tool calls with their results', () => {
const bundled: CoreMessage[] = [
{ role: 'user', content: 'Test' },
{
role: 'assistant',
content: [
{ type: 'tool-call', toolCallId: 'id1', toolName: 'tool1', args: {} },
{ type: 'tool-call', toolCallId: 'id2', toolName: 'tool2', args: {} },
],
},
{
role: 'tool',
content: [{ type: 'tool-result', toolCallId: 'id1', toolName: 'tool1', result: {} }],
},
{
role: 'tool',
content: [{ type: 'tool-result', toolCallId: 'id2', toolName: 'tool2', result: {} }],
},
];
const interleaved = properlyInterleaveMessages(bundled);
expect(interleaved).toHaveLength(5);
expect(validateArrayAccess(interleaved, 0, 'interleaved')?.role).toBe('user');
expect(validateArrayAccess(interleaved, 1, 'interleaved')?.role).toBe('assistant');
const interleaved1 = validateArrayAccess(interleaved, 1, 'interleaved');
const c1 = interleaved1?.content;
if (Array.isArray(c1) && c1[0] && typeof c1[0] === 'object' && 'toolCallId' in c1[0]) {
expect(c1[0].toolCallId).toBe('id1');
}
expect(validateArrayAccess(interleaved, 2, 'interleaved')?.role).toBe('tool');
const interleaved2 = validateArrayAccess(interleaved, 2, 'interleaved');
const c2 = interleaved2?.content;
if (Array.isArray(c2) && c2[0] && typeof c2[0] === 'object' && 'toolCallId' in c2[0]) {
expect(c2[0].toolCallId).toBe('id1');
}
expect(validateArrayAccess(interleaved, 3, 'interleaved')?.role).toBe('assistant');
const interleaved3 = validateArrayAccess(interleaved, 3, 'interleaved');
const c3 = interleaved3?.content;
if (Array.isArray(c3) && c3[0] && typeof c3[0] === 'object' && 'toolCallId' in c3[0]) {
expect(c3[0].toolCallId).toBe('id2');
}
expect(validateArrayAccess(interleaved, 4, 'interleaved')?.role).toBe('tool');
const interleaved4 = validateArrayAccess(interleaved, 4, 'interleaved');
const c4 = interleaved4?.content;
if (Array.isArray(c4) && c4[0] && typeof c4[0] === 'object' && 'toolCallId' in c4[0]) {
expect(c4[0].toolCallId).toBe('id2');
}
});
test('should handle mixed content (text + tool calls)', () => {
const mixed: CoreMessage[] = [
{
role: 'assistant',
content: [
{ type: 'text', text: 'Let me help you with that.' },
{ type: 'tool-call', toolCallId: 'id1', toolName: 'analyze', args: {} },
{ type: 'tool-call', toolCallId: 'id2', toolName: 'finalize', args: {} },
],
},
{
role: 'tool',
content: [{ type: 'tool-result', toolCallId: 'id1', toolName: 'analyze', result: {} }],
},
{
role: 'tool',
content: [{ type: 'tool-result', toolCallId: 'id2', toolName: 'finalize', result: {} }],
},
];
const interleaved = properlyInterleaveMessages(mixed);
expect(interleaved).toHaveLength(5);
expect(interleaved[0].role).toBe('assistant');
expect(interleaved[0].content).toEqual([
{ type: 'text', text: 'Let me help you with that.' },
]);
expect(interleaved[1].role).toBe('assistant');
const ic1 = interleaved[1].content;
if (Array.isArray(ic1) && ic1[0] && typeof ic1[0] === 'object' && 'toolCallId' in ic1[0]) {
expect(ic1[0].toolCallId).toBe('id1');
}
expect(interleaved[2].role).toBe('tool');
expect(interleaved[3].role).toBe('assistant');
const ic3 = interleaved[3].content;
if (Array.isArray(ic3) && ic3[0] && typeof ic3[0] === 'object' && 'toolCallId' in ic3[0]) {
expect(ic3[0].toolCallId).toBe('id2');
}
expect(interleaved[4].role).toBe('tool');
});
test('should handle conversation with follow-up questions', () => {
const conversation: CoreMessage[] = [
// First question
{ role: 'user', content: 'What is our revenue?' },
{
role: 'assistant',
content: [
{ type: 'tool-call', toolCallId: 't1', toolName: 'sql', args: { query: 'revenue' } },
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
toolCallId: 't1',
toolName: 'sql',
result: { revenue: 1000000 },
},
],
},
{
role: 'assistant',
content: 'Your revenue is $1M.',
},
// Follow-up question
{ role: 'user', content: 'What about profit?' },
{
role: 'assistant',
content: [
{ type: 'tool-call', toolCallId: 't2', toolName: 'sql', args: { query: 'profit' } },
],
},
{
role: 'tool',
content: [
{ type: 'tool-result', toolCallId: 't2', toolName: 'sql', result: { profit: 200000 } },
],
},
];
const result = properlyInterleaveMessages(conversation);
// Should remain mostly unchanged as it's already properly formatted
// (but IDs may be added to assistant messages with tool calls)
expect(result).toHaveLength(7);
expect(result[0]).toEqual(conversation[0]); // user message unchanged
expect(result[1].role).toBe('assistant');
expect(result[1].content).toEqual(conversation[1].content);
expect(result[2]).toEqual(conversation[2]); // tool result unchanged
expect(result[3]).toEqual(conversation[3]); // assistant text unchanged
expect(result[4]).toEqual(conversation[4]); // user message unchanged
expect(result[5].role).toBe('assistant');
expect(result[5].content).toEqual(conversation[5].content);
expect(result[6]).toEqual(conversation[6]); // tool result unchanged
});
});
describe('Real-world Conversation Pattern', () => {
test('should handle a complete conversation flow with multiple tool calls', () => {
const conversation: CoreMessage[] = [
{
role: 'user',
content: 'What are our top 5 products by revenue in the last quarter?',
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'sequentialThinking',
toolCallId: 'toolu_01Um5qidhhwormgMx9mASBv2',
args: {
thought: 'Analyzing the request for top 5 products...',
},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: { success: true },
toolName: 'sequentialThinking',
toolCallId: 'toolu_01Um5qidhhwormgMx9mASBv2',
},
],
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'submitThoughts',
toolCallId: 'toolu_01LA17JT7CUATsE4YX2Dy8oz',
args: {},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: {},
toolName: 'submitThoughts',
toolCallId: 'toolu_01LA17JT7CUATsE4YX2Dy8oz',
},
],
},
{
role: 'user',
content: 'Can you show me the year-over-year growth for these top products?',
},
{
role: 'assistant',
content: [
{
type: 'tool-call',
toolName: 'sequentialThinking',
toolCallId: 'toolu_01KkYSiZru6J8fvYdA9puoFX',
args: {
thought: 'Now analyzing year-over-year growth...',
},
},
],
},
{
role: 'tool',
content: [
{
type: 'tool-result',
result: { success: true },
toolName: 'sequentialThinking',
toolCallId: 'toolu_01KkYSiZru6J8fvYdA9puoFX',
},
],
},
];
// Verify the pattern is correct
expect(validateArrayAccess(conversation, 0, 'conversation')?.role).toBe('user');
expect(validateArrayAccess(conversation, 1, 'conversation')?.role).toBe('assistant');
const conv1 = validateArrayAccess(conversation, 1, 'conversation');
expect(conv1 ? isToolCallOnlyMessage(conv1) : false).toBe(true);
expect(validateArrayAccess(conversation, 2, 'conversation')?.role).toBe('tool');
expect(validateArrayAccess(conversation, 3, 'conversation')?.role).toBe('assistant');
const conv3 = validateArrayAccess(conversation, 3, 'conversation');
expect(conv3 ? isToolCallOnlyMessage(conv3) : false).toBe(true);
expect(validateArrayAccess(conversation, 4, 'conversation')?.role).toBe('tool');
expect(validateArrayAccess(conversation, 5, 'conversation')?.role).toBe('user');
expect(validateArrayAccess(conversation, 6, 'conversation')?.role).toBe('assistant');
expect(validateArrayAccess(conversation, 7, 'conversation')?.role).toBe('tool');
// Verify extraction preserves the structure (but may add IDs)
const extracted = extractMessageHistory(conversation);
expect(extracted).toHaveLength(8);
// Check the roles and structure are preserved
for (let i = 0; i < conversation.length; i++) {
const extractedItem = validateArrayAccess(extracted, i, 'extracted');
const conversationItem = validateArrayAccess(conversation, i, 'conversation');
expect(extractedItem?.role).toBe(conversationItem?.role);
expect(extractedItem?.content).toEqual(conversationItem?.content);
}
// Verify summary is correct
const summary = getConversationSummary(conversation);
expect(summary.userMessages).toBe(2);
expect(summary.assistantMessages).toBe(3);
expect(summary.toolCalls).toBe(3);
expect(summary.toolResults).toBe(3);
expect(summary.toolsUsed).toEqual(['sequentialThinking', 'submitThoughts']);
});
});
});