Skip to content

Exception for empty.tail #118

Description

@darthsuogles

Activity

  1. phi-dbq commented on Aug 9, 2017

    @phi-dbq
    Contributor

    Below is a small example in which the output tensor is scalar.

    vec_size = 17
    num_vecs = 41 * sc.defaultParallelism
    
    df = spark.createDataFrame([
        Row(idx=idx, vec=np.random.randn(vec_size).tolist())
        for idx in range(num_vecs)])
    
    analyzed_df = tfs.analyze(df)
    tfs.print_schema(analyzed_df)
    print(analyzed_df.count())
    
    with tf.Session():
        #x = tf.placeholder(tf.float64, shape=[None, vec_size])
        x = tfs.block(analyzed_df, 'vec')
        z = tf.reduce_mean(x)
        output_df = tfs.map_blocks([z], analyzed_df)
    

    TensorFrame is capable to infer the block size as

     |-- idx: long (nullable = true) long[41]
     |-- vec: array (nullable = true) double[41,17]
    

    But when trying to run output_df.show(), it throws the following error message.

    Py4JJavaError: An error occurred while calling o848.buildDF.
    : java.lang.UnsupportedOperationException: empty.tail
    	at scala.collection.TraversableLike$class.tail(TraversableLike.scala:421)
    	at scala.collection.mutable.ArrayOps$ofLong.scala$collection$IndexedSeqOptimized$$super$tail(ArrayOps.scala:246)
    	at scala.collection.IndexedSeqOptimized$class.tail(IndexedSeqOptimized.scala:129)
    	at scala.collection.mutable.ArrayOps$ofLong.tail(ArrayOps.scala:246)
    	at org.tensorframes.Shape.tail(Shape.scala:49)
    	at org.tensorframes.ColumnInformation$.structField(ColumnInformation.scala:82)
    	at org.tensorframes.impl.DebugRowOps$$anonfun$29.apply(DebugRowOps.scala:357)
    	at org.tensorframes.impl.DebugRowOps$$anonfun$29.apply(DebugRowOps.scala:351)
    	at scala.collection.immutable.Stream.map(Stream.scala:418)
    	at org.tensorframes.impl.DebugRowOps.mapBlocks(DebugRowOps.scala:351)
    	at org.tensorframes.impl.DebugRowOps.mapBlocks(DebugRowOps.scala:294)
    	at org.tensorframes.impl.PythonOpBuilder.buildDF(PythonInterface.scala:145)
    
  2. phi-dbq commented on Aug 9, 2017

    @phi-dbq
    Contributor

    It seems that tf.reduce_mean changes the leading (i.e. batch) dimension, which would lead to the final DataFrame not having the same number of rows as the input DataFrame.
    It will be nice to prompt the user that tensors in fetches must have the same leading dimension as the input tensors.

    For this example, changing tf.reduce_mean(x) to tf.reduce_mean(x, axis=1) gives the desired output.

  3. thunterdb commented on Aug 10, 2017

    @thunterdb
    Contributor

    @phi-dbq thanks. Yes, this error should be made more explicit. Just to confirm, in that case the shape of tf.reduce_mean(x) is going to be [17], if I am not mistaken?

  4. phi-dbq commented on Aug 10, 2017

    @phi-dbq
    Contributor

    @thunterdb thanks. Currently tf.reduce_mean without specifying the axis parameter will reduce along all dimensions. It became a scalar.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions