Pure Java/Scala code for writing Tensorflow TFRecords data file

I'm trying to write a pure Java/Scala implementation of the Tensorflow RecordWriter class in order to convert Spark DataFrame into TFRecords file. According to the documentation, in TFRecords, each record is formated as follow:

uint64 length
uint32 masked_crc32_of_length
byte   data[length]
uint32 masked_crc32_of_data

And the CRC mask

masked_crc = ((crc >> 15) | (crc << 17)) + 0xa282ead8ul

Currently, I compute the CRC with guava implementation with the following code:

import com.google.common.hash.Hashing

object CRC32 {
  val kMaskDelta = 0xa282ead8

  def hash(in: Array[Byte]): Int = {
    val hashing = Hashing.crc32c()

  def mask(crc: Int): Int ={
    ((crc >> 15) | (crc << 17)) + kMaskDelta

The rest of my code is:

The data encoding part is done with the following piece of code:

  object LittleEndianEncoding {
   def encodeLong(in: Long): Array[Byte] = {
    val baos = new ByteArrayOutputStream()
    val out = new LittleEndianDataOutputStream(baos)

  def encodeInt(in: Int): Array[Byte] = {
    val baos = new ByteArrayOutputStream()
    val out = new LittleEndianDataOutputStream(baos)


The record are generated with protocol buffer:

import com.google.protobuf.ByteString
import org.tensorflow.example._

import collection.JavaConversions._
import collection.mutable._

object TFRecord {

  def int64Feature(in: Long): Feature = {

    val valueBuilder = Int64List.newBuilder()


  def floatFeature(in: Float): Feature = {
    val valueBuilder = FloatList.newBuilder()

  def floatVectorFeature(in: Array[Float]): Feature = {
    val valueBuilder = FloatList.newBuilder()


  def bytesFeature(in: Array[Byte]): Feature = {
    val valueBuilder = BytesList.newBuilder()

  def makeFeatures(features: HashMap[String, Feature]): Features = {

  def makeExample(features: Features): Example = {


And here is an example of how I use things together in order to generate my TFRecords file:

val label = TFRecord.int64Feature(1)
val feature = TFRecord.floatVectorFeature(Array[Float](1, 2, 3, 4))
val features = TFRecord.makeFeatures(HashMap[String, Feature]  ("feature"->feature, "label"-> label))
val ex = TFRecord.makeExample(features)
val exSerialized = ex.toByteArray()
val length = LittleEndianEncoding.encodeLong(exSerialized.length)
val crcLength =  LittleEndianEncoding.encodeInt(CRC32.mask(CRC32.hash(length)))
val crcEx = LittleEndianEncoding.encodeInt(CRC32.mask(CRC32.hash(exSerialized)))

val out = new FileOutputStream(new File("test.tfrecords"))

When I try to read the file I got inside Tensorflow with TFRecordReader, I get the following error:

W tensorflow/core/common_runtime/executor.cc:1076] 0x24cc430 Compute status: Data loss: corrupted record at 0

I suspect that the CRC mask computation is not correct or the endianness between java and c++ generated file are not the same.

1 Answers

FWIW, the Tensorflow team has provided utility code for reading/writing TFRecords, which can be found in the ecosystem repo

