추석 준비를 AI에 물어보는 분이 늘었습니다. 그런데 같은 질문을 넣어도 서비스마다 답이 다릅니다.
이번 주에 나온 조사 두 건이 그 차이를 보여줍니다. SR타임스는 챗GPT·클로드·네이버 AI탭·업스테이지 솔라 네 곳에 같은 질문을 넣어 답을 비교했고, 이투데이는 에이아이빅스랩이 챗GPT·제미나이·클로드·퍼플렉시티 네 곳의 추석 선물 답변을 분석한 결과를 전했습니다.
두 조사를 겹쳐 놓으면 한 가지가 눈에 띕니다. 네 곳 중 챗GPT만 양쪽에서 다른 쪽에 서 있습니다. 우연으로 보이지만 원인은 하나로 모입니다.
■ 두 조사가 같은 곳에서 어긋났다
어디서 갈렸을까요?
먼저 귀성길입니다. 서울에서 부산까지 정체를 피할 출발 시간을 묻자 클로드와 솔라, 네이버 AI탭은 모두 티맵의 추석 교통 예측 자료를 근거로 9월 23일 늦은 밤부터 24일 새벽 사이를 꼽았습니다. 클로드는 25일 낮 12시에 출발하면 7시간 3분이 걸릴 수 있다는 수치까지 댔습니다.
챗GPT만 달랐습니다. 올해 추석 특별교통대책의 시간대별 전망을 충분히 확인하기 어렵다고 먼저 밝힌 뒤, 평년 명절 패턴을 기준으로 24일 오전 4~5시를 제시했습니다. 예상 이동시간은 5시간에서 6시간 30분으로 폭이 넓었습니다.
선물 쪽 조사에서도 같은 일이 벌어졌습니다. 제미나이와 클로드, 퍼플렉시티에서는 스팸이 1위, 정관장이 2위였는데 챗GPT에서만 이마트가 1위였습니다. 챗GPT가 상품 브랜드를 적게 대신 구매처를 자주 언급한 결과라는 분석이 붙었습니다.
한쪽만 보면 그냥 다른 답입니다. 두 조사를 겹쳐 보면 챗GPT가 두 번 다 '구체적인 최신 값' 대신 '일반적인 틀'을 내놓은 쪽에 서 있습니다. 그래서 성능 차이가 아니라 정보를 어디서 가져오느냐의 차이로 읽는 편이 맞습니다. 성능 차가 아니라 출처 차.
■ 학습 시점은 6월, 추석은 9월
첫 번째 원인은 학습 시점입니다. AI 모델은 어느 시점까지의 데이터로 학습을 끝내고, 그 뒤 일은 모릅니다.
앤트로픽이 공개한 모델별 학습 시점은 이렇습니다. 클로드 오퍼스 5.5와 페이블 5.1은 2026년 6월, 오퍼스 5는 2026년 5월, 소네트 5와 페이블 5는 2026년 1월까지입니다. 헤이쿠 4.5는 2025년 7월, 오퍼스 3은 2023년 8월까지입니다. 앤트로픽은 이 시점 이후의 일은 모델이 모를 수 있고, 더 최근 일을 물으면 정확하지 않은 정보를 낼 수 있다고 적어 뒀습니다.
여기서 중요한 건 가장 최신 모델도 6월까지라는 점입니다. 올해 추석 교통 예측이나 이번 달 선물세트 가격은 어느 모델의 학습 데이터에도 들어 있지 않습니다. 학습은 6월, 연휴는 9월.
그러니 최신 값이 답에 들어왔다면 학습한 내용이 아니라 그 자리에서 검색해 온 것입니다. 검색에 붙지 않았다면 남는 건 평년 패턴뿐입니다. 챗GPT가 귀성길에서 내놓은 답이 정확히 그 모양이었습니다.
■ 검색에 붙었는지 보는 법
그럼 내가 쓰는 AI가 지금 검색에 붙어 있는지는 어떻게 알까요?
두 번째 원인은 그때 검색이 돌았느냐입니다. 이건 이용자가 확인할 수 있습니다.
챗GPT는 최신 정보가 도움이 될 질문이면 알아서 웹을 검색한다고 오픈AI가 설명합니다. 다만 알아서 판단하는 것이라 안 돌 수도 있습니다. 직접 시키려면 입력창에 슬래시(/)를 치고 Search를 고르면 됩니다. 이미 받은 답을 다시 만들고 싶으면 새로고침 표시를 눌러 '웹 검색'을 고르는 방법도 있습니다. 웹 검색은 무료 요금제에서도 쓸 수 있고, 로그인하지 않아도 됩니다.
클로드는 채팅창 왼쪽 아래 '+' 버튼을 눌러 'Web search'를 켜면 체크 표시가 붙습니다. 다시 누르면 꺼집니다. 다만 새로 바뀐 클로드 화면에는 이 토글이 아예 없고, 도움이 될 때 알아서 검색한다고 앤트로픽은 설명합니다.
켜졌는지 확인하는 가장 확실한 신호는 출처입니다. 앤트로픽은 검색이 돈 답변에는 항상 인용이 붙어 확인할 수 있다고 적었고, 오픈AI도 검색을 쓴 답변에는 인용이 포함될 수 있으며 인용을 누르면 원문이 열린다고 안내합니다. 출처 표시가 하나도 없는 답이라면 그 숫자는 검색해 온 값이 아닐 가능성이 큽니다. 인용 없는 최신 수치는 일단 보류.
출처가 붙었다고 끝이 아닙니다. 오픈AI는 검색 결과와 인용이 불완전하거나 오래됐거나 틀릴 수 있다고 못박고, 인용한 원문을 열어 답을 실제로 뒷받침하는지 보고 언제 게시되거나 갱신됐는지 확인하라고 권합니다. 정확성이 중요한 일이면 권위 있는 출처를 쓰라는 말도 함께 붙어 있습니다.
■ 같은 질문인데 나만 다른 답을 받는 이유
조사와 똑같이 물었는데 내 화면 답만 다르다면 어떻게 봐야 할까요?
세 번째 원인은 이용자마다 조건이 다르다는 것입니다. 조사에서 나온 답과 내 화면의 답이 달라지는 자리이기도 합니다.
위치가 먼저입니다. 오픈AI는 챗GPT가 IP 주소를 바탕으로 대략적인 위치를 써서 지역 결과를 낸다고 설명합니다. 기기 위치 공유는 선택이고 기본은 꺼져 있습니다. 그래서 귀성길이나 주변 가게를 물을 때는 질문에 도시나 동네, 우편번호를 직접 넣으라고 권합니다. VPN을 쓰면 IP로 추정한 위치가 달라질 수 있다는 설명도 있습니다.
메모리도 영향을 줍니다. 오픈AI는 메모리가 켜져 있으면 저장된 내용을 검색어를 다시 쓰는 데 쓸 수 있다고 밝혔습니다. 채식이고 샌프란시스코에 산다고 말해 둔 적이 있으면 식당을 물을 때 '샌프란시스코 채식 식당'으로 검색할 수 있다는 예를 들었습니다. 지난 대화에 남긴 취향이 이번 추석 답에도 따라붙는다는 뜻입니다.
SR타임스가 비교를 하면서 매번 새 대화창에 같은 질문을 넣고 무료 서비스와 기본 설정을 기준으로 삼았다고 밝힌 것도 이 때문입니다. 조건을 맞추지 않으면 비교 자체가 성립하지 않습니다.
■ 연휴에 물을 때 확인할 것
정리하면 이렇습니다.
- 날짜가 걸린 질문은 검색을 직접 켭니다 — 교통, 영업시간, 시세처럼 이번 주에만 맞는 값은 학습 데이터에 없습니다
- 답에 출처가 붙었는지 봅니다 — 인용이 하나도 없으면 검색해 온 값이 아닐 수 있습니다
- 출처를 눌러 게시일을 봅니다 — 오픈AI가 직접 권하는 확인 절차입니다
- 위치는 질문에 적어 넣습니다 — 기기 위치 공유는 기본이 꺼져 있습니다
- 여러 AI에 물을 거면 조건을 맞춥니다 — 새 대화창, 같은 문장, 같은 설정이어야 비교가 됩니다
- 브랜드 순위는 인기 순위가 아닙니다 — 조사를 한 곳도 답변에 언급된 양상일 뿐 판매량이나 선호도를 뜻하지 않는다고 못박았습니다
마지막 항목은 특히 선물 고를 때 걸립니다. 이투데이가 전한 조사에서 전체 답변의 34.7%에는 브랜드가 한 번도 나오지 않았습니다. AI가 한우나 과일, 한과처럼 품목만 말하고 제품명은 대지 않은 경우입니다. 이름이 자주 나온 브랜드가 좋은 선물이라는 뜻은 아닙니다.
AI에 추석을 맡기기 어렵다는 말은 아닙니다. 다만 답이 갈리는 자리가 성능이 아니라 정보를 어디서 가져왔느냐에 있다는 걸 알면, 어떤 답을 그대로 믿고 어떤 답을 다시 확인할지 고르기 쉬워집니다.
출처 · SR타임스 — [추석기획] AI에 맡긴 추석…귀성길부터 차례상까지 답은 ‘제각각’ · 이투데이 — AI가 추천한 추석선물 2위는 ‘정관장’…1위는? · OpenAI 공식 도움말 — Searching the web with ChatGPT · Anthropic 공식 도움말 — How up-to-date is Claude's training data? · Anthropic 공식 도움말 — Enable and use web search · Anthropic 공식 도움말 — When should I use web search, extended thinking, and research?
More people are asking AI to help with Chuseok preparations. Yet the same question gets different answers depending on the service.
Two surveys published this week show the gap. SR Times put identical questions to four services - ChatGPT, Claude, Naver AI Tab and Upstage Solar - and compared the answers. Etoday reported an analysis by AIBIX Lab of Chuseok gift answers from ChatGPT, Gemini, Claude and Perplexity.
Lay the two surveys side by side and one thing stands out: of the four services, only ChatGPT lands on the other side in both. It looks like coincidence, but the cause is single.
■ Both surveys diverged at the same point
Where exactly did they split?
Start with the drive home. Asked when to leave Seoul for Busan to avoid congestion, Claude, Solar and Naver AI Tab all cited TMAP's Chuseok traffic forecast and picked late on September 23 through the early hours of the 24th. Claude went as far as a figure: leaving at noon on the 25th could take 7 hours 3 minutes.
ChatGPT alone differed. It first said it could not sufficiently verify this year's hour-by-hour holiday traffic plan, then offered 4-5am on the 24th based on typical holiday patterns, with a wide estimate of five to six and a half hours.
The gift survey shows the same pattern. Gemini, Claude and Perplexity all put Spam first and Jung Kwan Jang second; only ChatGPT put E-Mart first. The analysis attributed this to ChatGPT naming fewer product brands and mentioning places to buy more often.
Seen alone, each is just a different answer. Seen together, ChatGPT twice sat on the side of the general framework rather than the specific current figure. That reads as a difference in where the information came from, not in capability.
■ Training stops in June; Chuseok is in September
The first cause is the knowledge cutoff. A model finishes training on data up to some date and does not know what came after.
Anthropic publishes the dates. Claude Opus 5.5 and Fable 5.1 were trained on data up until June 2026, Opus 5 up until May 2026, and Sonnet 5 and Fable 5 up until January 2026. Haiku 4.5 goes to July 2025 and Opus 3 to August 2023. Anthropic notes these models may not be aware of events after their cutoff, and may give inaccurate information about more recent events.
The point is that even the newest model stops at June. This year's holiday traffic forecast and this month's gift-set prices are in no model's training data.
So if a current figure appears in an answer, it was searched for on the spot rather than recalled. Without a search, all that remains is the typical pattern - which is exactly the shape of the answer ChatGPT gave on the drive home.
■ How to tell whether search actually ran
So how do you tell whether your AI is actually connected to search right now?
The second cause is whether search ran at all, and this is something you can check.
OpenAI says ChatGPT may search the web automatically when a question would benefit from current information - which also means it may not. To force it, type / in the composer and select Search. You can also regenerate an existing answer with the refresh control and choose to search the web. Web search is available on the free plan, and even without signing in.
In Claude, click the "+" button at the lower left of the chat window and select Web search; a checkmark appears, and clicking again turns it off. Anthropic notes that the newer Claude experience has no toggle at all - it searches when that helps.
The most reliable signal is citations. Anthropic says every response from a search includes citations so you can verify them, and OpenAI says responses using web search may include citations that open their source when selected. An answer with no sources at all probably did not fetch its numbers.
Citations alone are not enough. OpenAI states plainly that search results and citations can be incomplete, outdated or incorrect, and advises opening a cited source to check it supports the answer, reviewing when it was published or updated, and using an authoritative source when accuracy matters.
■ Why your answer differs from the survey's
And if you asked exactly what the survey asked but got a different answer?
The third cause is that conditions differ per user - which is also where your screen departs from what a survey reported.
Location comes first. OpenAI says ChatGPT may use an approximate location based on your IP address for local results, while device location sharing is optional and off by default. That is why it advises including your city, neighborhood or postal code in the question itself. A VPN can change the location inferred from your IP.
Memory matters too. OpenAI says that if memory is enabled, ChatGPT may use saved memories when rewriting a search query - its own example is someone who mentioned being vegan and living in San Francisco getting a search for vegan restaurants there. Preferences left in earlier chats follow you into this year's Chuseok answers.
That is also why SR Times noted it entered each question in a fresh conversation and used free services on default settings. Without matching conditions, the comparison does not hold.
■ What to check when you ask over the holiday
In short:
- Turn search on yourself for anything date-bound - traffic, opening hours and prices that hold only this week are not in the training data
- Check whether the answer carries sources - no citations at all may mean nothing was fetched
- Open a source and check its date - this is the step OpenAI itself recommends
- Write the location into the question - device location sharing is off by default
- If you ask several services, match the conditions - fresh chat, same wording, same settings
- A brand ranking is not a popularity ranking - the firm behind the survey said its index reflects how brands appeared in answers, not sales or consumer preference
That last one bites hardest when choosing gifts. In the survey Etoday reported, 34.7% of all answers named no brand at all - the AI listed categories like beef, fruit or traditional sweets without naming a product. A brand appearing often does not make it a good gift.
None of this means AI is a poor holiday assistant. But knowing that answers diverge over where the information came from, rather than over capability, makes it much easier to decide which answers to take as they are and which to check again.
Sources · SR Times · Etoday · OpenAI Help Center · Anthropic Help Center · Anthropic Help Center · Anthropic Help Center