Welcome to OStack Knowledge Sharing Community for programmer and developer-Open, Learning and Share
Welcome To Ask or Share your Answers For Others

Categories

0 votes
342 views
in Technique[技术] by (71.8m points)

hadoop - Hive every Insert query creates a new file in Hdfs file system

On every insert query one files gets created with 000000_0_copy* in Hdfs file system.

Is this the default behaviour of hive and Hdfs ?

Is there any concept of compaction if yes then How does the comapaction work?

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…
Welcome To Ask or Share your Answers For Others

1 Answer

0 votes
by (71.8m points)

HDFS is an append only filesystem, meaning to modify (UPDATE/DELETE statements) any portion of a file that is already written, one must rewrite the entire file and replace the old file, or write a new file to insert even a single record.

Compaction isn't an automatic process. You need to write your own code to query one table, then insert into another format like parquet/orc


与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…
Welcome to OStack Knowledge Sharing Community for programmer and developer-Open, Learning and Share
Click Here to Ask a Question

...