Kimi K3 underperforms US frontier AI models in planning cyberattacks

แหล่งที่มา Cryptopolitan

A joint evaluation report has stated that Moonshot AI’s Kimi K3 is not up to par with the strongest American AI systems when it comes to building software exploits and running simulated network attacks. 

The United Kingdom’s AI Security Institute (AISI) and the US Center of AI Standards and Innovation (CAISI) ran the evaluation, and they made their findings public on July 23.

Moonshot is expected to open the model’s full weights to the public on July 27, and the findings say that Kimi K3’s safeguards did nothing to stop it from helping with offensive cyber work.

What did the joint evaluation discover about Kimi K3?

The two institutes ran Kimi K3 through ExploitBench, a Carnegie Mellon test that scores how far a model can push an exploit through to completion. It is based on 41 vulnerabilities found in Chrome’s V8 engine after 2023. Kimi K3 scored 32%. The leading models from the US averaged 76.2%, and China’s GLM-5.2 came in at 24%.

Arbitrary code execution, which hands an attacker full control of a target machine, is the benchmark’s most dangerous outcome. US models reached it on 20 of the 41 tasks on average. Kimi K3 reached it on none.

GLM-5.2 also failed to get there. So while Kimi K3 now leads open-weight rivals in cyber capability, it comes short at the point where a real attack would do the most damage.

Halfway through a 32-step network breach

The second test, called “The Last Ones,” drops a model into a simulated corporate network with about 20 hosts spread across four subnets. 

It would take a human specialist around 20 hours to work the full 32-step path. Kimi K3 averaged step 17, while the most capable US models averaged 28.5, and GLM-5.2 managed 11.

However, the evaluators noted that Kimi K3 finished the whole chain once in ten tries and stayed under the 100-million-token ceiling the evaluators set. 

The best US models solved it six or seven times out of ten. The institutes read that lone success as evidence that the model has the capability; however, it cannot summon it on demand. They noted the range has no active defenders and a built-in path to the target, so it flatters any attacker.

Does Kimi K3 have safeguards that enable it to say no?

The evaluators pointed out that the model’s willingness to carry out tasks without pushback was a major risk. They stated, “Kimi K3 is capable of autonomously attacking small, weakly defended, and vulnerable enterprise systems when directed to do so and given initial network access.” 

AISI has previously warned that rising open-model capability creates “a persistent and irreversible risk of misuse.” Once Kimi K3’s weights are public on July 27, its behavior can no longer be gated by whoever hosts it.

Why do Kimi K3 cyber scores lag when compared to its general scores?

Before this report, Kimi K3 had strong general benchmarks, and that has not changed; however, it highlighted its shortcomings when its cyber capabilities were put under test. The Chinese 2.8-trillion-parameter model reportedly beat Claude 4.8 and GPT-5.5 and also topped some leaderboards.

According to the institutes, this may be possible as a result of how the model may have been built.

White House science director Michael Kratsios accused Moonshot on July 22 of distilling Anthropic’s Fable model, using its outputs as training data. 

Anthropic’s classifiers block advanced offensive cyber prompts. This means that a training set that is scraped from Claude responses would not offer enough material on those exact skills. 

Kimi K3 does not have those classifiers, and this allowed it to match Western models on general work while it still came short in exploit ability.

If you're reading this, you’re already ahead. Stay there with our newsletter.

ข้อจำกัดความรับผิดชอบ: เพื่อการอ้างอิงเท่านั้น ผลการดำเนินงานในอดีตไม่ได้บ่งบอกถึงผลลัพธ์ในอนาคต
placeholder
แนวโน้มราคาน้ำมันดิบ: ความตึงเครียดในตะวันออกกลางผลักดันราคาน้ำมันให้สูงขึ้น ราคาจะยังสามารถปรับตัวสูงขึ้นได้อีกหรือไม่หลังจากทะลุระดับ $100? ผลกระทบจากการทวีความรุนแรงอย่างต่อเนื่องของความตึงเครียดทางภูมิรัฐศาสตร์ในตะวันออกกลาง ส่งผลให้น้ำมันดิบเบรนท์ ( UKOIL) กลับมาทะลุเหนือระดับจิตวิทยาที่ 100 ดอลลาร์ต่อบาร์เรลอีกครั้งในรอบส
ผู้เขียน  TradingKey
12 ชั่วโมงที่แล้ว
ผลกระทบจากการทวีความรุนแรงอย่างต่อเนื่องของความตึงเครียดทางภูมิรัฐศาสตร์ในตะวันออกกลาง ส่งผลให้น้ำมันดิบเบรนท์ ( UKOIL) กลับมาทะลุเหนือระดับจิตวิทยาที่ 100 ดอลลาร์ต่อบาร์เรลอีกครั้งในรอบส
placeholder
ราคาทองคําปรับตัวลดลงต่อเนื่องเนื่องจากความกังวลเงินเฟ้อหนุนการเก็งการขึ้นดอกเบี้ยของเฟดและค่าเงินดอลลาร์สหรัฐท่ามกลางมาตรการทองคํา (XAU/USD) ดึงดูดผู้ขายเป็นวันที่สองติดต่อกันในวันศุกร์และอ่อนค่าลงต่อเนื่องต่ำกว่าระดับราคา $4,050 ในช่วงเซสชันเอเชีย
ผู้เขียน  FXStreet
13 ชั่วโมงที่แล้ว
ทองคํา (XAU/USD) ดึงดูดผู้ขายเป็นวันที่สองติดต่อกันในวันศุกร์และอ่อนค่าลงต่อเนื่องต่ำกว่าระดับราคา $4,050 ในช่วงเซสชันเอเชีย
placeholder
ข่าวหุ้นวันนี้ น้ำมันแตะ 100 ดอลลาร์กดตลาด หุ้นเทคโดนขาย และ SET ต้องเลือกหุ้นรับแรงกระแทกทันทุกกระแสการเงิน สรุปข่าวเด่น Forex หุ้น ทองคำ คริปโตฯ และเศรษฐกิจรอบวัน วิเคราะห์แนวโน้มตลาดด้วยข้อมูลเชิงลึก อ่านง่าย เข้าใจไว อัปเดตล่าสุดที่นี่
ผู้เขียน  Mitrade
16 ชั่วโมงที่แล้ว
ทันทุกกระแสการเงิน สรุปข่าวเด่น Forex หุ้น ทองคำ คริปโตฯ และเศรษฐกิจรอบวัน วิเคราะห์แนวโน้มตลาดด้วยข้อมูลเชิงลึก อ่านง่าย เข้าใจไว อัปเดตล่าสุดที่นี่
placeholder
สรุปภาวะตลาดวันนี้: ราคาน้ำมันทะลุ 100 ดอลลาร์, กระตุ้นความกังวลด้านเงินเฟ้อ, ขณะที่งบลงทุน (Capex) ด้าน AI เผชิญการตรวจสอบ และการดิ่งลง 14% ของเทสลา (Tesla) ฉุดกลุ่มเทคโนโลยีเกาะติดแนวโน้มตลาดTradingKey - ตลาดเผชิญกับผลกระทบซ้ำสองจากราคาน้ำมันที่พุ่งสูงขึ้นและความกังขาเกี่ยวกับผลตอบแทนจากการลงทุนใน AI ซึ่งถูกจุดชนวนจากการเพิ่มขึ้นของค่าใช้จ่ายฝ่ายทุนของกูเกิล (GOOGL) และเ
ผู้เขียน  TradingKey
19 ชั่วโมงที่แล้ว
เกาะติดแนวโน้มตลาดTradingKey - ตลาดเผชิญกับผลกระทบซ้ำสองจากราคาน้ำมันที่พุ่งสูงขึ้นและความกังขาเกี่ยวกับผลตอบแทนจากการลงทุนใน AI ซึ่งถูกจุดชนวนจากการเพิ่มขึ้นของค่าใช้จ่ายฝ่ายทุนของกูเกิล (GOOGL) และเ
placeholder
ทองคำยังคงเหนือระดับ 4,100 ดอลลาร์ ขณะที่ค่าเงินดอลลาร์สหรัฐฯ อ่อนแอชดเชยการเก็งการขึ้นดอกเบี้ยของเฟดท่ามกลางความตึงเครียดระหว่างสหรัฐฯ กับอิหร่านทองคํา (XAUUSD) ยืนเหนือระดับ $4,100 ในช่วงตลาดเอเชียวันพฤหัสบดี และในขณะนี้ดูเหมือนว่าจะหยุดการย่อตัวเล็กน้อยจากระดับสูงสุดในรอบมากกว่าสองสัปดาห์เมื่อวันก่อนหน้าไว้ได้
ผู้เขียน  FXStreet
เมื่อวาน 07: 15
ทองคํา (XAUUSD) ยืนเหนือระดับ $4,100 ในช่วงตลาดเอเชียวันพฤหัสบดี และในขณะนี้ดูเหมือนว่าจะหยุดการย่อตัวเล็กน้อยจากระดับสูงสุดในรอบมากกว่าสองสัปดาห์เมื่อวันก่อนหน้าไว้ได้
goTop
quote